Retriever resources
The Claude Code & AI glossary
111 terms you keep seeing in AI threads, on Claude Code tutorials, in docs. Plain definitions, what they mean in practice, and where you actually run into them.
A
A2A (Agent2Agent)An open protocol that lets AI agents from different vendors talk to each other.
A2A is a protocol Google introduced in 2025 so that agents built by different companies, on different frameworks, can discover each other, hand off tasks and exchange results. Think of it as a shared language between agents.
It complements MCP rather than replacing it: MCP connects one agent to tools and data, A2A connects agents to other agents.
AgentAn LLM plus tools, instructions, and a loop that lets it act on its own.
In AI, an agent is a model that can take action on its own to reach a goal. It does not just answer a question — it uses tools (search the web, run a calculation, call an API, read a file), checks what happened, and decides the next step. It works in a loop until the task is done or you stop it.
In Claude Code specifically, the agent is a Claude instance running in your terminal: it can read and edit files in your project, run shell commands, search the web, open pull requests. A sub-agent is the same idea at smaller scale — the main agent spawns one to handle a scoped job and gets just the answer back.
Agent SDKAnthropic's library for building your own agents on the same engine as Claude Code.
The Claude Agent SDK gives developers the same tools, agent loop and context management that power Claude Code, as a Python or TypeScript library. Instead of chatting with Claude Code in a terminal, you embed that agent inside your own product or script: it can read files, run commands, search the web, call MCP servers and spawn sub-agents, with hooks and permissions.
Use it when you want a custom agent that runs in your own process. If you would rather have Anthropic host the agent for you, see Managed Agents.
Agent teamsSeveral Claude Code sessions working together, coordinated by a lead session.
An agent team is a group of full Claude Code sessions. One session is the lead: it splits the work, assigns tasks and synthesizes the results. The others, the teammates, each work in their own context window, share a task list and can message each other directly instead of only reporting back to the lead.
That is the difference with sub-agents, which just return a result to whoever spawned them. Teams suit work where agents should challenge each other, like a code review from three angles or a bug hunt with competing hypotheses. They cost more tokens, and in Claude Code they are still experimental and off by default.
Agentic loopThe cycle an agent repeats: think, act with a tool, look at the result, decide the next step.
Every agent runs on the same basic loop. The model reads the task, picks an action (search, read a file, run a command), sees what came back, and decides whether it is done or what to try next. It keeps going until the goal is reached or it hits a limit. This loop is what separates an agent from a chatbot that answers once and stops.
In Claude Code, each pass through the loop shows up in the transcript as a tool call followed by its result. The harness runs the loop; the model decides what happens inside it.
AGENTS.mdA shared instructions file for AI coding agents, read by Codex, Cursor and Claude Code.
AGENTS.md is a plain Markdown file at the root of a repository that tells coding agents how the project works: how to run it, which conventions to follow, which traps to avoid. It plays the same role as CLAUDE.md, but it is a cross-tool convention, so one file serves every agent your team uses.
Claude Code reads AGENTS.md when the project has no CLAUDE.md. To use both, a common setup is a CLAUDE.md that imports it with a single line: `@AGENTS.md`.
AnthropicThe company that builds Claude.
Anthropic is the AI lab that builds Claude. Founded in 2021 in San Francisco by ex-OpenAI researchers, they train and ship the Claude family of models (Opus, Sonnet, Haiku), the API that lets developers call them, and Claude Code — the terminal tool that turned Claude into a coding agent. They publish a lot about AI safety, which shows up in how Claude refuses certain requests and how the harness is designed.
AntigravityGoogle's agentic coding IDE — its answer to Cursor and Claude Code.
Antigravity is Google's coding IDE built around autonomous AI agents that plan and execute multi-step coding tasks. It runs Gemini models and integrates with the wider Google Cloud and AI Studio stack. Announced in late 2025.
Direct competitor to Cursor and Claude Code, on the Google side.
APIApplication Programming Interface — the door a program uses to talk to another service.
An API is how two pieces of software talk to each other. Instead of going through a website, your code sends a request to a URL and gets a structured reply back, usually in JSON. Almost every modern app is glued together by APIs: when you log in with Google, pay with Stripe, or load a Slack message, that is an API call happening in the background. To use one you typically need an API key — a long secret string the provider gives you in your account dashboard.
For AI products, "calling the API" means sending a prompt to the company's servers and receiving the model's reply. Every lab ships one (OpenAI, Anthropic, Google, Mistral). The Claude API is what sits underneath claude.ai, Claude Code, and every third-party tool built on top of Claude.
ArtifactA live web page Claude builds from your work and publishes to a private link you can share.
Some results are easier to look at than to read: a dashboard, an annotated diff, a side-by-side comparison of options. An artifact is a self-contained interactive page Claude builds from your conversation and publishes on claude.ai. It stays private until you share it, and Claude can update it in place as the work continues.
From Claude Code, you just ask for one ("make an artifact that..."). The page can even pull fresh data from your connectors each time someone opens it.
Auto modeA permission mode where a second model reviews Claude's actions instead of asking you each time.
Without it, a coding agent keeps stopping to ask "can I run this command?". In auto mode, a separate classifier model checks each action in the background: routine ones go through, risky ones get blocked. You stop babysitting the agent without handing it a blank cheque.
In Claude Code, it is one of the permission modes you cycle through with Shift+Tab, and the default starting mode on Pro, Max and Team plans. It is what lets long tasks, like a /goal, run unattended.
Auto-compactWhen a long conversation is summarised to fit back into the context window.
Every LLM has a limit on how much text it can read at once (the context window). When a conversation grows close to that limit, products usually deal with it one of two ways: drop the oldest messages, or summarise them. Compaction is the second approach — older parts of the chat are condensed into a short summary so newer content still fits.
In Claude Code, this happens automatically: a notice appears in the terminal when it kicks in. Some details may be lost, but the session continues without crashing into the limit.
B
Background taskA long-running command that runs separately so it does not block your main work.
A background task is a long-running command — a build that takes two minutes, a dev server that runs forever, a test suite, a data sync — started in a way that does not block what you are doing. Very common in development: you launch a dev server in one terminal and keep coding in another.
In Claude Code, the agent can start a background task itself: "Command running in background" appears in the transcript, and when the task finishes Claude reads its output. Useful for letting tests run while the agent keeps editing files.
BashThe default shell on Mac and Linux, used to run command-line programs.
Bash is the program that interprets the commands you type in a terminal on Mac or Linux: `ls` lists files in a folder, `cd folder` moves into a folder, `git commit` saves a change. It is the most common shell in the developer world.
When you ask Claude Code (or any coding agent) to "run the tests" or "install this package", it uses a Bash tool under the hood: it types the command into your shell the same way you would. You see exactly what ran in the transcript, and you can approve or block specific commands.
Batch APIBulk API mode: upload many requests at once, get all answers back within 24 hours, half the price.
Most AI APIs return each call within seconds (synchronous). A batch API is the opposite: you upload a big file of requests — say, 10,000 prompts — wait up to 24 hours, and receive all the answers back at once. The trade is speed for cost: you wait longer but pay roughly half the price. It is the right choice for offline jobs where latency does not matter — classifying a lead list, enriching a database, summarising thousands of documents.
OpenAI, Anthropic, and most major providers offer one.
C
Cache (prompt caching)A trick to reuse parts of a long prompt across calls so they cost less and run faster.
Every time you call an LLM, the model re-reads your full prompt from scratch, which costs tokens (and money). Prompt caching is a shortcut offered by most providers: if the start of your prompt is identical to a previous request — a system prompt, a long document, an attached PDF — it gets stored for a few minutes and re-used instead of re-processed. Cached tokens cost about a tenth of normal tokens and process faster.
Claude Code uses this automatically (with CLAUDE.md, the system prompt, files you keep re-reading), which is why long sessions do not get exponentially expensive.
Chain of thoughtWhen the model writes out its reasoning step by step before answering.
If you ask an LLM "what is 17 × 24?", it does better when it works through the steps (17 × 20 = 340, then 17 × 4 = 68, total 408) than when it guesses straight to a number. That step-by-step working is called chain of thought, and it improves accuracy on hard problems. Used in every modern model.
Claude does this on its own when a problem needs it — you will often see Claude write a short plan or list its assumptions before acting. Extended thinking is the same idea turned up: more tokens spent reasoning before replying, for genuinely hard problems.
ChannelsA way to push messages from Telegram, Discord or a webhook into a running Claude Code session.
Normally you talk to Claude Code by typing in the terminal. A channel lets outside events land in the session you already have open: a message you send from your phone, a failed CI build, an alert from your monitoring. Claude reads it, does the work on your machine, and can reply through the same channel.
Channels are plugins you install and switch on per session. Telegram, Discord and iMessage are supported. The feature is in research preview.
ChatGPTOpenAI's chat app — the product that put LLMs in mainstream hands in late 2022.
ChatGPT is the consumer chat interface from OpenAI, launched in November 2022. It is a web and mobile app where you chat with one of OpenAI's GPT models. Underneath it calls the OpenAI API, which is what most third-party apps use directly. Its launch is the reason most people first heard of LLMs at all.
On the Anthropic side, the equivalent is claude.ai.
CheckpointA saved snapshot of your project state that you can roll back to.
A checkpoint is any moment in your project's history you can return to later. The best checkpoints, by far, are Git commits: each commit is a snapshot `git reset` or `git checkout` can put you back at. Best practice when working with any AI agent: commit before letting it do something big, so if it goes sideways you are one command away from the original state.
Claude Code also saves a checkpoint automatically before each prompt you send. Run `/rewind` (or press Esc twice) to roll back the code, the conversation, or both. It only tracks edits Claude made with its own file tools, not changes made by shell commands, so Git stays your real safety net.
Claude in ChromeThe browser extension that lets Claude drive your Chrome: open tabs, click, type, read pages.
Claude in Chrome connects Claude to your actual browser. It opens tabs, clicks, fills forms and reads what is on the page, including console errors, using the sites you are already logged into. When it hits a login page or a CAPTCHA, it stops and hands over to you.
From Claude Code, launch with `claude --chrome` (or run `/chrome`) and ask for things like "open localhost:3000 and check the signup form" or "pull the prices from this page into a CSV".
CLAUDE.mdA file in your repo where you write instructions Claude reads at the start of every session.
Put project conventions, stack notes, deploy steps, and "always do X" rules here. Claude treats it as durable context.
CLICommand Line Interface, a program you talk to by typing commands in a terminal.
Claude Code is a CLI: you launch it with `claude` in your terminal and chat with it there.
CodebaseThe full set of source files for a project.
When Claude Code "reads the codebase" it grep-searches and reads relevant files on demand, not the whole project at once.
CodexOpenAI's coding agent. Also the name of an older OpenAI code model — check the date.
OpenAI uses the name Codex for two different things. In 2021 it was the code model that powered the first version of GitHub Copilot. Then in 2025 OpenAI brought the name back for a new coding agent that runs in your terminal and in the cloud, edits your repo, and runs tests.
The 2025 Codex is the direct OpenAI equivalent of Claude Code.
CommitA versioned snapshot of a change in Git.
A commit records what changed, who changed it, and a message describing why. Claude can create commits for you with `git commit`.
Compact (/compact)The command that summarizes the conversation so far to free up context.
Long sessions fill the context window with old messages, file contents and logs. `/compact` replaces that history with a summary, so Claude keeps the essentials and gets room to work again. You can add a hint about what to keep, like `/compact focus on the API changes`.
It is the manual version of Auto-compact. Its cousin `/clear` goes further: it starts a fresh conversation, and only your project memory (CLAUDE.md) carries over.
Computer useA capability that lets Claude control your screen, mouse, and keyboard.
Used for tasks no API can do, like clicking through a native app. Risky: scope what Claude can touch, and watch what it does.
Context engineeringThe craft of choosing what information ends up in the model's context window.
A bigger skill than "prompt engineering". Includes deciding what to retrieve, summarise, cache, and exclude so the model has just enough to do the job.
Context rotThe drop in answer quality when a model's context gets too long or too cluttered.
A bigger context window does not mean the model uses all of it equally well. As a conversation piles up irrelevant logs, failed attempts and contradictory instructions, models start missing details, repeating mistakes or forgetting early instructions. That gradual decay is called context rot.
The fixes are what context engineering is about: start a fresh session for a new task, compact long ones, and hand noisy research to sub-agents so only their conclusion comes back.
Context windowHow much text the model can read in one go, measured in tokens.
Claude Opus and Sonnet support up to 1 million tokens, roughly 750 pages of text. Past the window, older content gets compacted or dropped.
CoworkThe tab in the Claude desktop app for longer agentic work beyond coding.
The Claude desktop app has three tabs: Chat, Cowork and Code. Cowork is where you hand Claude longer tasks and let it work through them: research, documents, spreadsheets. It hosts Dispatch, a persistent conversation where you drop tasks and check back later.
When a task turns out to be development work, like fixing a bug or opening a pull request, Cowork can start a Claude Code session to handle it.
CursorA code editor (a fork of VS Code) built around AI assistance.
Cursor is one option for editing with AI. Claude Code itself runs in a terminal and can be paired with Cursor, VS Code, or any editor.
D
DeepSeekA Chinese open-weights LLM family known for strong reasoning at very low cost.
DeepSeek is a Chinese AI lab that publishes open-weights models (DeepSeek-V3, DeepSeek-R1). "Open weights" means anyone can download and run the model on their own hardware, unlike Claude or GPT which are only available through an API.
Their late-2024 and early-2025 releases caught attention by matching frontier models at a fraction of the training cost, and forced US labs to rethink their pricing.
DiffThe list of lines added and removed between two versions of a file.
Every edit Claude proposes is shown as a diff so you can see exactly what changed before accepting.
E
Edit toolThe built-in Claude Code action that modifies an existing file in place.
Replaces an exact string with another. Safer than rewriting the whole file because the change is small and reviewable.
EffortThe setting that controls how hard the model thinks before it answers, from low to max.
Recent models decide on their own how much to reason at each step. The effort level sets the budget: low is fast and cheap for simple tasks, high and above give deeper reasoning on hard problems but take longer and use more tokens.
In Claude Code, change it with `/effort` (low, medium, high, xhigh, max). To get deeper thinking on a single message without changing the setting, see Ultrathink.
EmbeddingA vector representation of text used for semantic search and similarity.
Two pieces of text with similar meaning have nearby vectors. Used in RAG and recommendation systems.
EndpointA URL that an API listens on for requests.
The Claude API has endpoints for messages, batches, files, and so on.
EvalsTests that measure whether a model, a prompt or an agent actually does its job well.
Short for evaluations. An eval is a set of test cases plus a way to score the output: did the agent find the right answer, follow the format, avoid the forbidden action? You run it before and after a change to know whether you improved things or just moved the problem.
For anyone building with AI, evals replace "it looked good on the three examples I tried". They can be scored by code, by people, or by another model (see LLM-as-a-judge).
Extended thinkingA mode where the model spends extra tokens reasoning before it replies.
Trades latency for quality on hard problems like long refactors or proofs. The thinking is visible to the user and to following turns.
On recent Claude models, thinking is adaptive: the model decides how much to think at each step, and the effort level is the main dial. See Effort and Ultrathink.
F
Fast modeA Claude Code setting that makes Opus answer up to 2.5x faster, at a higher price.
Fast mode runs the same Opus model with a configuration that favors speed over cost: same quality, shorter wait, higher price per token. It fits live debugging and quick back-and-forth, less so long unattended tasks where speed does not matter.
Toggle it with `/fast`. It is in research preview, and on subscription plans it is billed through usage credits.
G
GeminiGoogle's flagship LLM family — the equivalent of GPT (OpenAI) and Claude (Anthropic).
Gemini is the model family from Google DeepMind, first released late 2023. It powers Google's Gemini chat (the product), NotebookLM, Antigravity, and any "powered by Gemini" product. Tiers go from Ultra to Pro to Flash, ordered from most capable to fastest.
Known for very long context windows (1M+ tokens) and strong multimodal abilities — it reads images, audio, and video natively.
GitA version control system that tracks every change to your code.
Claude Code uses Git constantly: to read diffs, create commits, push branches, open pull requests.
GitHubA hosting service for Git repositories with reviews, issues, and CI.
Claude can interact with GitHub through the `gh` CLI or the GitHub MCP server.
Goal (/goal)A Claude Code command: set a finish line and Claude keeps working until it is reached.
Normally Claude stops after each turn and waits for you. With `/goal` you state a condition, like `/goal all tests in test/auth pass and lint is clean`, and Claude keeps working turn after turn without you prompting each step. After every turn, a separate small model checks whether the condition holds. The goal ends when it is met, when the checker judges it impossible, or when you run `/goal clear`.
Write the condition as something Claude can prove: a test result, a build that passes, an empty queue. Pair it with auto mode so it runs unattended.
GPTOpenAI's model family — GPT-3, GPT-4, GPT-4o, GPT-5, o1, o3, and others.
GPT stands for Generative Pre-trained Transformer. It is the name of the model family that powers ChatGPT and the OpenAI API. Versions go GPT-3 (2020), GPT-4 (2023), GPT-4o, GPT-5, and so on. The "o" variants (o1, o3) are reasoning-focused models tuned to spend more compute thinking before answering.
The exact version you reach depends on the product or the API model ID, which changes every few months.
GrokThe LLM from xAI (Elon Musk's AI company), built into X / Twitter.
Grok is the model family from xAI. It powers the Grok chat on X (formerly Twitter) and is available via API. Known for its access to live X data — posts, replies, trends — and a less-restricted tone than ChatGPT or Claude.
Directly competes with the other frontier model families (GPT, Claude, Gemini).
GuardrailsThe limits that keep an AI agent from doing things it should not.
Guardrails are the rules and checks around an agent: which tools it can use, which files it can touch, which actions need a human yes, what it must refuse. Some live in the prompt as instructions; the stronger ones are enforced by the system around the model, because a model can ignore an instruction but not a blocked permission.
In Claude Code, guardrails take the form of permission modes, allow and deny rules in settings.json, hooks that block risky commands, and the sandbox.
H
HaikuThe smallest and fastest Claude model in the lineup.
Good for high-volume cheap work like classification, light extraction, or quick UI calls.
HallucinationWhen the model confidently invents a fact, API, or file that does not exist.
The fix is grounding: have the model read the real code or docs first instead of guessing.
HarnessThe technical envelope around the model that handles tools, permissions, memory, and the loop.
Claude Code is a harness around a Claude model. The same Claude model behaves very differently inside Claude Code, the API, or claude.ai because the harness around it is different.
You can build on the same harness yourself: the Agent SDK runs it inside your own code, and Managed Agents runs it on Anthropic's servers.
Headless modeRunning Claude Code without the chat interface, from a script or an automation.
Headless (or non-interactive) mode means you give the agent a task in one command, it runs to completion and exits, with nobody typing in between. That is how you plug an agent into automations: a nightly job, a CI check, a script that processes a folder of files.
In Claude Code it is the `-p` flag: `claude -p "summarize the changes in this PR"`. Add `--output-format json` to get a result other programs can read.
HermesAn open-source LLM family fine-tuned by Nous Research, usually on top of Llama.
Hermes is a family of open-source models (Hermes 2, Hermes 3, and others) fine-tuned by Nous Research. Where Llama is the raw open-weights model from Meta, Hermes is a community fine-tune aiming for stronger reasoning, instruction-following, and tool use.
You run it yourself on local hardware or call it through an inference provider — there is no single official API like Claude or GPT.
HookA shell command the harness runs automatically when a specific event happens.
Examples: run a linter after every edit, block certain commands, notify a Slack channel when Claude finishes. Configured in `settings.json`.
Human-in-the-loopA setup where a person approves or corrects the agent at key moments.
Full autonomy is not always the goal. Human-in-the-loop means the agent does the work but stops at chosen checkpoints for a person to validate: before sending an email, deleting data, spending money or shipping code. The skill is picking the right checkpoints: too many and you are babysitting, too few and mistakes go out unseen.
In Claude Code, permission prompts and plan mode are the built-in ways to keep a human in the loop.
I
IDEIntegrated Development Environment, an editor enriched for coding (Cursor, VS Code, JetBrains).
Claude Code can run inside an IDE's terminal and share context with the editor (open file, selection, diagnostics).
InferenceA single call to the model: prompt in, completion out.
Latency and cost are measured per inference. Caching and batching reduce both.
J
JSONA simple text format for structured data, built from keys, values, lists, and nested objects.
Used everywhere: tool inputs and outputs, API payloads, config files. Claude can produce strict JSON on demand.
L
LatencyHow long a request takes from send to response.
First-token latency matters for chat UIs; total latency matters for automations. Caching, smaller models, and streaming all reduce perceived latency.
LlamaMeta's open-weights LLM family — the foundation for most open-source AI work.
Llama (Llama 2, Llama 3, Llama 4) is Meta's model family, published with open weights. Anyone can download and run them on their own hardware, unlike Claude or GPT which are only available through an API. It is the most-used base for community fine-tunes — Hermes, Code Llama, and dozens of others.
When people say "I run an open-source LLM", they usually mean a Llama-derived model.
LLMLarge Language Model, the kind of neural network that powers Claude.
Trained on text to predict the next token. With instruction tuning and tools, it becomes a working assistant.
LLM-as-a-judgeUsing one model to grade the output of another.
Checking thousands of AI answers by hand does not scale. So you give a second model the answer and a grading rubric (is it correct, on topic, well sourced?) and let it score. It is the most common way to run evals on open-ended tasks where there is no single right answer.
The judge can be wrong too, so good setups spot-check its verdicts against human judgment. Claude Code's /goal runs on this idea: a small model judges after each turn whether your condition is met.
Local modelAn LLM that runs entirely on your own machine, no network call.
Better for privacy and offline work, weaker than frontier models. Not what Claude Code uses by default.
Long-running agentAn agent that works on a task for hours, not seconds.
Early AI tools answered in one shot. Long-running agents take on work that needs many steps over a long stretch: migrate a codebase, research a market, clear a backlog. They need things short tasks do not: a memory of what they already did, a way to recover from errors, and a clear definition of done.
In Claude Code, /goal, background tasks, workflows and routines are the pieces that make long runs possible. On the API side, Managed Agents are built for them.
Loop (/loop)A Claude Code command that reruns a prompt on a schedule while your session is open.
Some tasks are about checking back: is the deploy done, did CI pass, any new review comments? `/loop 5m check the deploy` reruns that prompt every five minutes. Leave out the interval and Claude picks the pace itself, checking often while things move and less when they are quiet.
It only runs while the session is open, and recurring loops expire after seven days. For work that should run with your laptop closed, use a Routine.
M
Managed AgentsAgents hosted and run by Anthropic, set up through the Claude API.
Building an agent usually means building the loop, the tool execution, the sandbox and the storage yourself. Claude Managed Agents provides all of that as a service: you define the agent (model, instructions, tools, skills), and Anthropic runs it in a secure cloud sandbox where it can read files, run commands and browse the web, and keeps its session history.
It is built for long-running and asynchronous work, and it is in beta. For an agent that runs inside your own process instead, see Agent SDK.
MCPModel Context Protocol, an open standard for plugging external tools and data into AI assistants.
An MCP server exposes resources (read-only data) and tools (actions). Claude Code can connect to many at once: GitHub, Notion, Linear, your database, internal APIs.
MemoryA persistent store the assistant can write notes into across conversations.
In Claude Code, memory lives in a folder of Markdown files. Useful for user preferences, project context, and feedback that should carry over.
MistralA French AI lab known for open-weights models and a strong European positioning.
Mistral AI is a Paris-based lab founded in 2023, the highest-profile European challenger to OpenAI and Anthropic. They publish open-weights models (Mistral 7B, Mixtral, Codestral) alongside commercial APIs and a chat product called Le Chat.
The brand is the French / EU bet on sovereign AI: independent from US providers, open to local hosting.
ModelA specific trained network you can call by name (e.g. claude-opus-4-7).
The Claude lineup is Opus, Sonnet, Haiku, ordered from most capable to fastest. Each has versioned releases.
MultimodalA model that can read more than just text, typically images and sometimes audio or video.
Claude can read screenshots, PDFs, and diagrams. Useful for "look at this UI and tell me what is wrong".
O
OpenAIThe AI lab behind ChatGPT and the GPT family.
OpenAI is the San Francisco AI lab founded in 2015. It built and ships ChatGPT, the GPT model family, the OpenAI API, and a range of other models for images, video, and coding. The launch of ChatGPT in November 2022 triggered the wave of AI products you see today.
Direct competitors: Anthropic (Claude) and Google DeepMind (Gemini).
OpusThe most capable model in the Claude lineup.
Used for hard reasoning, long agentic runs, and large refactors. Slower and more expensive than Sonnet or Haiku.
OrchestratorThe agent that splits a big task and hands the pieces to other agents.
In a multi-agent system, one agent acts as the manager. It breaks the task down, sends each piece to a specialized agent (a researcher, a coder, a reviewer), then collects and merges what comes back. The workers stay focused on their small job; the orchestrator keeps the big picture.
In Claude Code, the main session orchestrates when it spawns sub-agents, a lead session orchestrates an agent team, and a workflow moves the orchestration itself into a script.
Output styleA Claude Code setting that changes the tone and format of every answer in a session.
Instead of asking "be shorter" in every prompt, you pick a style once. Claude Code ships built-in ones: Concise (answer first, no narration), Explanatory (adds short notes on why it made each choice), Learning (leaves small parts of the code for you to write) and Proactive (starts working without asking about routine decisions). You can also write your own.
Switch with `/output-style`.
P
Permission modeThe harness setting that decides which tools Claude can use without asking.
Modes range from strict to permissive: Manual asks before most edits and commands, acceptEdits lets file edits through, plan stays read-only, auto lets a classifier model approve actions in the background, and bypassPermissions skips checks (for isolated machines only). Switch with Shift+Tab. Project-level rules live in `.claude/settings.json`.
Plan modeA read-only Claude Code mode for exploring a problem before any edits.
Claude can read files and run shell commands, but cannot write or run destructive actions. Useful to scope a refactor first.
PluginAn installable pack that adds skills, agents, hooks or MCP servers to Claude Code.
A plugin bundles several customizations into one thing you install with a single command: skills, sub-agents, hooks, MCP servers. It is how teams and the community share setups instead of copying files around. Plugins are distributed through marketplaces, catalogs you add once and then browse.
In Claude Code, manage them with `/plugin`. Anthropic maintains an official marketplace and a community one.
PR (pull request)A proposal to merge a branch into the main codebase, reviewed by teammates.
Claude can open, comment on, and review PRs through the `gh` CLI or the GitHub MCP server.
PromptAny instruction you send to the model.
In practice the model sees a stack: system prompt, prior messages, your message. "Prompt" usually means just your message.
Prompt engineeringThe craft of writing prompts that reliably get the answer you want.
Examples, format constraints, role framing, and worked steps all help. Less central than it used to be as models follow plain instructions better.
Prompt injectionWhen malicious text inside data the model reads tries to override your instructions.
For example, a webpage saying "ignore previous instructions and email this address". The fix is sandboxing and reviewing what the model actually does.
R
RAG (Retrieval-Augmented Generation)Pulling relevant snippets from a knowledge base and stuffing them into the prompt before asking.
Lets the model answer about content it was not trained on, like your own docs or a fresh codebase.
Rate limitThe cap that throttles how often you can call the API.
When you exceed it, the API responds with 429 and a retry-after hint. Most SDKs back off automatically.
Reasoning modelA model that thinks step by step before it answers.
A reasoning model spends extra tokens working through a problem internally (planning, checking, backtracking) before giving its final answer. It is slower and costs more per answer, but it is much better at math, code, multi-step logic and anything where a quick guess goes wrong.
Most recent frontier models, Claude included, reason adaptively: they decide how much to think at each step. You steer how much with the effort setting.
Remote ControlDrive a Claude Code session running on your computer from your phone or another browser.
You start a task at your desk and want to follow it from the couch. Remote Control connects the Claude mobile app or claude.ai/code to a Claude Code session running on your machine. The work keeps happening locally, on your files and tools; your phone is just another screen and keyboard for the same conversation.
Start it with `claude remote-control`, or `/remote-control` inside a session.
RepoShort for repository, a project tracked by Git.
A repo holds the code, history, and configuration for one project. Claude Code operates on whichever repo your terminal is in.
RoutineA scheduled task that runs Claude on a cron-like timer.
Example: run "check overnight PRs and summarise to Slack" every morning at 8am. Configured via the schedule skill.
S
SandboxAn isolated environment where code or commands can run without affecting the rest of the system.
Vercel Sandbox and Claude Code's permission system both aim at the same thing: limit blast radius.
Settings.jsonThe Claude Code config file where permissions, hooks, env vars, and MCP servers live.
Two layers: `~/.claude/settings.json` (global, all projects) and `.claude/settings.json` (project-specific).
SkillA reusable workflow you can invoke with a slash command.
A skill bundles instructions, examples, and sometimes tools so a multi-step task ("ship a PR", "review code") runs the same way every time.
Slash commandA shortcut typed in chat as `/name` that triggers a skill or built-in action.
Examples in Claude Code: `/help`, `/clear`, `/review`. Custom slash commands are how teams package shared workflows.
SonnetThe mid-tier Claude model: nearly as capable as Opus, much faster and cheaper.
The default workhorse for most coding and agentic tasks.
Status lineThe customizable bar at the bottom of Claude Code.
The status line shows useful information under the prompt at a glance: which model you are on, how full the context window is, what the session has cost, which git branch you are on. You decide what it shows by pointing it at a small script.
Set it up with `/statusline`, or in settings.json.
Structured outputsForcing a model to answer in an exact data format, like a JSON template.
When a program reads the model's answer, free text is a problem. With structured outputs you give the model a schema, the list of fields and types you expect, and the answer is guaranteed to match it. No more broken automations because the model added a sentence before the JSON.
It is a feature of the Claude API, and `claude -p --json-schema` brings the same idea to Claude Code scripts.
Sub-agentA separate Claude instance spawned by the main agent for a scoped task.
Used for parallel research, isolating long context, or running risky steps in a fresh sandbox. The main agent gets just the result back.
System promptThe base instructions the harness sends to the model on every request, invisible in normal use.
Sets identity, tools, rules, and tone. Different products (Claude Code, claude.ai, API) ship different system prompts on the same underlying model.
T
TerminalThe text-based interface where you type commands to your computer.
On Mac: Terminal.app or iTerm. Claude Code lives here.
ThinkingSee Extended thinking.
TokenThe unit the model counts in: roughly 4 characters or 3/4 of a word in English.
Pricing and context windows are measured in tokens. A 1M-token window holds about 750,000 words.
Tool useWhen the model decides to call a tool (file read, shell command, API) instead of just writing text.
Each tool has a name, a schema, and a description. The model picks one, fills in the inputs, and the harness runs it and returns the result.
TurnOne message in a conversation, either from the user or the assistant.
A multi-turn run is a back-and-forth. Each turn adds tokens to the context window.
U
UltrathinkA keyword that asks Claude Code for deeper reasoning on one message.
Type ultrathink anywhere in your prompt and Claude Code asks the model to reason harder on that turn only, without changing your effort setting for the rest of the session. Handy for the one tricky question in an otherwise routine session.
Other phrases like "think hard" are treated as normal text. To change reasoning for the whole session, use /effort.
User promptThe message you actually send the model, distinct from the system prompt.
In chat UIs this is what you type into the box. Everything else around it (system prompt, prior turns, tool definitions) is set by the harness.
V
Vibe codingLetting the AI write the code while you describe intent and review the result.
Coined by Andrej Karpathy. Works for prototypes; demands real review for anything that ships to production.
VS CodeVisual Studio Code, a free editor from Microsoft, the most widely used IDE today.
Claude Code integrates with VS Code via an extension that shares context with the editor.
W
WebhookA URL another service calls when an event happens (PR opened, deploy finished, message received).
How chat bots and CI integrations receive notifications. The Vercel Chat SDK and the GitHub MCP both speak webhooks.
WorkflowA script that runs many Claude sub-agents in parallel on one big task.
Some jobs are too big for one conversation: audit every file in a repo, migrate 500 components, cross-check a research question across dozens of sources. A workflow is a script Claude writes that launches many sub-agents, usually in phases, and gathers their results while your session stays free.
In Claude Code, ask for one in plain words or with the keyword ultracode, run the built-in `/deep-research`, or set `/effort ultracode` to let Claude plan workflows on its own. A run can use many more tokens than a normal conversation.
WorktreeA separate working directory pointing at a different branch of the same Git repo.
Lets you experiment on a branch without touching your main checkout. Claude Code can spawn an agent inside a worktree so its changes stay isolated.
Y
YAMLA human-friendly format for config files and frontmatter.
Skills and memory files in Claude Code use YAML frontmatter at the top to declare metadata like name, description, and type.