The short answer
MCP costs tokens twice. In the menu: every tool definition is loaded before your first word. And in the round trips: each tool call is another pass over the whole conversation, and every result goes back through the model. Tool Search shrinks the menu. Only code removes the round trips.
In this guide
How do you measure what MCP costs you?
Before you change anything, look at where your tokens go. In Claude Code, three commands show it:
/contextshows what is using your context window, MCP tools included./mcplists your servers and lets you disable the ones you don’t use./usageon a subscription, breaks your recent plan usage down by skills, subagents, plugins and individual MCP servers.
Two numbers matter: how big the tool menu is, and how many tool calls a typical task makes. The first is in /context. The second you can count in the transcript: every tool call is another request to the model.
Why does MCP use so many tokens? Two taxes
The menu tax
To pick a tool, the model has to read its definition. MCP servers publish all of theirs, so a big setup starts the conversation already full. Anthropic measured five servers with 58 tools at about 55,000 tokens before the first message; GitHub’s 35 tools alone came to about 26,000. A long menu also hurts accuracy: past 30–50 tools, Anthropic says, Claude gets worse at picking the right one.
This tax is already being cut at the source. Claude Code defers MCP tool definitions by default: only names and server instructions enter the context until a tool is used. Cursor does the same with dynamic context discovery, which cut tokens by 46.9% in its own test of runs that called MCP tools. Anthropic’s Tool Search Tool reports 85% fewer tokens.
The round-trip tax
A tool call is not a function call inside the model. It is a new request: the model writes the call, the runtime runs it, the result comes back, and the model reads the whole conversation again to write the next call. Claude Code’s own documentation says it plainly: it sends your full conversation with every request, and each time Claude uses tools it sends another request carrying that batch of tool results. Prompt caching makes the re-read cheaper, not free.
Results pay a second time. Anthropic’s example: moving a meeting transcript from Google Drive to Salesforce can mean processing an additional 50,000 tokens, because the text flows through the model on its way.
Tool by tool
One program
Tool Search does nothing about any of this. On our own test, fixing a release on a server with 152 tools, Claude Code made 26 model calls with Tool Search on.
What reduces MCP token usage? The fixes, compared
Each fix solves one tax or both. None is free, and the right one depends on how you use MCP.
| Fix | Cuts the menu | Cuts the round trips | Effort | The catch |
|---|---|---|---|---|
| Disconnect the servers you don’t use | PartlyFor those servers | No | Minutes | You lose those tools until you reconnect them. |
| Tool Search (deferred tool definitions) | YesBuilt into Claude Code and Cursor | No | None, on by default | Each lookup is one more step, and the calls themselves are unchanged. |
| A CLI instead of an MCP server (gh, aws…) | YesNo per-tool listing | NoStill one command per call | Only where a good CLI exists | You give up MCP’s reach: many services have an MCP server and no CLI. |
| An MCP gateway | YesUsually | PartlyOnly with a code mode | A service to run and secure | Built for teams and governance, not for one laptop. |
| Let the model write code (code mode), built by you | YesA few tools instead of hundreds | YesOne program per task | High: a secure sandbox, limits and monitoring | You build and maintain the environment, and Anthropic warns it adds operational overhead. |
| Delta MCP | YesThree tools for every server | YesOne program per task | Install it: two minutes, free | Built for tasks with several steps. |
Code mode is the pattern Anthropic describes in Code execution with MCP and Cloudflare in Code Mode. Delta MCP is that pattern, ready-made, for the servers you already use.
What does Delta MCP change?
Delta MCP is a free app that sits between your AI apps and their MCP servers. Your AI sees three tools instead of every tool of every server, and writes each task as one typed program. Delta MCP checks the program before the first call, makes the fewest calls, applies them all or nothing, reads the result back and keeps it in Activity, ready to undo.
That is code mode without the work: the checks, the all-or-nothing apply and the undo come with the app, on your computer. Your servers stay exactly as they are, and every original tool stays callable by name.
How many tokens does Delta MCP save?
Claude Code with Tool Search on, Sonnet 5.5, two to three runs per variant, the end result checked in each service.
| Task | Tokens | Model calls | Measured |
|---|---|---|---|
| Fix a release, as codeCerberus, 152 tools | 13.4× fewer | 26 → 6.5 | 5 October |
| Add tests, as textCerberus, 152 tools | 15.8× fewer | 16.6 → 4 | 2 October |
| Clean up 20 late ticketsA ticket manager | 4.8× fewer | 5 → 2 | 2 October |
| Log in, add 3 contacts, list themMicrosoft Playwright MCP | 3.9× fewer | 12.6 → 4 | 2 October |
| A documentation questionContext7 | 1.2× fewer | 4 → 3 | 5 October |
The more steps a task has, the more Delta MCP saves.
All measurements and the method
Do I still need Delta MCP if I use Tool Search?
Yes, when your tasks have several steps. Tool Search keeps the menu small, but every step is still a pass. Delta MCP removes those passes: your AI writes the task as one program. Our measurements were taken with Tool Search on, so the savings come on top of it.
Questions
Why does MCP use so many tokens?
Two reasons. The model reads every tool definition before it can choose one, and every tool call is a separate request that re-reads the whole conversation, with each result flowing back through the model. The first is the menu tax, the second the round-trip tax.
How many tokens do MCP tools use?
It depends on your servers. Anthropic measured five servers with 58 tools at about 55,000 tokens before the first message, and GitHub’s 35 tools alone at about 26,000. Claude Code’s /context command shows your own number.
Does Tool Search fix MCP token usage?
It fixes the menu, not the round trips. Anthropic reports 85% fewer tokens for the menu, and Cursor’s equivalent cut tokens by 46.9% in runs that called MCP tools. But each tool call is still its own request over the whole conversation.
Is code execution with MCP really cheaper?
On multi-step tasks, often. Anthropic’s example fell from 150,000 to 2,000 tokens, and Cloudflare exposed more than 2,500 API endpoints in about 1,000 tokens. Both come with a cost: you need a safe place to run the code. In our own measurements Delta MCP used 3.3–24.1× fewer tokens on multi-step tasks.
Does Delta MCP save tokens?
Yes, on tasks with several steps: 3.3–24.1× fewer tokens in our tests, because your AI writes one program instead of one call per step.
Do I have to change my MCP servers?
No. Delta MCP uses them exactly as they are, and every original tool stays callable by name: 152 of 152 on the largest server we tried.
How do I check how many tokens my MCP servers use?
In Claude Code, /context shows what fills your context window, MCP tool definitions included. /mcp lists your servers and lets you switch off the ones you don’t use, and on a subscription /usage breaks recent usage down by individual MCP server.
Sources
- Anthropic, Code execution with MCP: Building more efficient agents (4 Nov 2025) The two problems with direct tool calls; the 150,000 to 2,000 tokens example; the 50,000-token transcript example; the sandbox warning.
- Anthropic, Advanced tool use (24 Nov 2025) 58 tools on five servers at about 55,000 tokens; 85% fewer tokens with the Tool Search Tool.
- Anthropic, Tool search tool (documentation) Claude’s accuracy at picking a tool degrades beyond 30–50 tools.
- Claude Code, Manage costs effectively (documentation) MCP tool definitions deferred by default; /context, /mcp and /usage; the full conversation sent with every request; CLIs and MCP overhead.
- Cursor, Dynamic context discovery (6 Jan 2026) 46.9% fewer tokens in runs that called an MCP tool, in Cursor’s own A/B test.
- Cloudflare, Code Mode: give agents an entire API in 1,000 tokens (20 Feb 2026) More than 2,500 endpoints in about 1,000 tokens, against 1.17 million as a plain MCP server.
- Delta MCP, Benchmarks Every measurement of Delta MCP on this page, with its method.
Keep reading
- Claude Code usage limit reached? What your MCP servers costWhat Claude Code's "You've hit your limit" means (session, weekly, Opus, spend), how /usage shows what your MCP servers cost, and what to cut first.
- MCP vs CLI for AI agents: where the token bill comes fromMCP or CLI for your AI agent? What each costs in tokens, what the measurements show, when a CLI wins, when MCP does, and where code beats both.
- MCP config file locations: Claude, Cursor, VS Code and moreWhere Claude Desktop, Claude Code, Cursor, VS Code, Windsurf, Gemini CLI and Codex keep their MCP config: file path, JSON key, remote servers and tool limits.