<!-- How to reduce MCP token usage, and why MCP burns tokens · https://www.mcpdelta.com/guides/reduce-mcp-token-usage · updated 2026-10-07 -->

> How to reduce MCP token usage: measure it with /context, see the two costs (tool menu and round trips), and compare every fix, Tool Search included.

[Delta MCP](https://www.mcpdelta.com/) [Guides](https://www.mcpdelta.com/guides)

Guide

# How to reduce MCP token usage, and why MCP burns tokens

Updated 7 October 20267 min read By the Delta MCP team

The short answer

MCP costs tokens twice. In the menu: every tool definition is loaded before your first word. And in the round trips: each tool call is another pass over the whole conversation, and every result goes back through the model. Tool Search shrinks the menu. Only code removes the round trips.

## How do you measure what MCP costs you?

Before you change anything, look at where your tokens go. In Claude Code, three commands show it:

-   `/context` shows what is using your context window, MCP tools included.
-   `/mcp` lists your servers and lets you disable the ones you don’t use.
-   `/usage` on a subscription, breaks your recent plan usage down by skills, subagents, plugins and individual MCP servers.

Two numbers matter: how big the tool menu is, and how many tool calls a typical task makes. The first is in `/context`. The second you can count in the transcript: every tool call is another request to the model.

## Why does MCP use so many tokens? Two taxes

### The menu tax

To pick a tool, the model has to read its definition. MCP servers publish all of theirs, so a big setup starts the conversation already full. Anthropic measured five servers with 58 tools at about 55,000 tokens before the first message; GitHub’s 35 tools alone came to about 26,000. A long menu also hurts accuracy: past 30–50 tools, Anthropic says, Claude gets worse at picking the right one.

This tax is already being cut at the source. Claude Code defers MCP tool definitions by default: only names and server instructions enter the context until a tool is used. Cursor does the same with dynamic context discovery, which cut tokens by 46.9% in its own test of runs that called MCP tools. Anthropic’s Tool Search Tool reports 85% fewer tokens.

### The round-trip tax

A tool call is not a function call inside the model. It is a new request: the model writes the call, the runtime runs it, the result comes back, and the model reads the whole conversation again to write the next call. Claude Code’s own documentation says it plainly: it sends your full conversation with every request, and each time Claude uses tools it sends another request carrying that batch of tool results. Prompt caching makes the re-read cheaper, not free.

Results pay a second time. Anthropic’s example: moving a meeting transcript from Google Drive to Salesforce can mean processing an additional 50,000 tokens, because the text flows through the model on its way.

Tool by tool

One program

Illustration. Each bar is one pass over the conversation. Tool Search makes the menu segment shorter; it does not make the staircase shorter. In our measured task, 26 model calls became 6.5.

Tool Search does nothing about any of this. On our own test, fixing a release on a server with 152 tools, Claude Code made 26 model calls with Tool Search on.

## What reduces MCP token usage? The fixes, compared

Each fix solves one tax or both. None is free, and the right one depends on how you use MCP.

What reduces MCP token usage? The fixes, compared
| Fix | Cuts the menu | Cuts the round trips | Effort | The catch |
| --- | --- | --- | --- | --- |
| Disconnect the servers you don’t use | Partly For those servers | No | Minutes | You lose those tools until you reconnect them. |
| Tool Search (deferred tool definitions) | Yes Built into Claude Code and Cursor | No | None, on by default | Each lookup is one more step, and the calls themselves are unchanged. |
| A CLI instead of an MCP server (gh, aws…) | Yes No per-tool listing | No Still one command per call | Only where a good CLI exists | You give up MCP’s reach: many services have an MCP server and no CLI. |
| An MCP gateway | Yes Usually | Partly Only with a code mode | A service to run and secure | Built for teams and governance, not for one laptop. |
| Let the model write code (code mode), built by you | Yes A few tools instead of hundreds | Yes One program per task | High: a secure sandbox, limits and monitoring | You build and maintain the environment, and Anthropic warns it adds operational overhead. |
| Delta MCP | Yes Three tools for every server | Yes One program per task | Install it: two minutes, free | Built for tasks with several steps. |

Code mode is the pattern Anthropic describes in _Code execution with MCP_ and Cloudflare in _Code Mode_. Delta MCP is that pattern, ready-made, for the servers you already use.

## What does Delta MCP change?

Delta MCP is a free app that sits between your AI apps and their MCP servers. Your AI sees three tools instead of every tool of every server, and writes each task as one typed program. Delta MCP checks the program before the first call, makes the fewest calls, applies them all or nothing, reads the result back and keeps it in Activity, ready to undo.

That is code mode without the work: the checks, the all-or-nothing apply and the undo come with the app, on your computer. Your servers stay exactly as they are, and every original tool stays callable by name.

[How it works, in detail](https://www.mcpdelta.com/how-it-works)

## How many tokens does Delta MCP save?

Claude Code with Tool Search on, Sonnet 5.5, two to three runs per variant, the end result checked in each service.

What we measured
| Task | Tokens | Model calls | Measured |
| --- | --- | --- | --- |
| Fix a release, as code Cerberus, 152 tools | **13.4× fewer** | 26 → 6.5 | 5 October |
| Add tests, as text Cerberus, 152 tools | **15.8× fewer** | 16.6 → 4 | 2 October |
| Clean up 20 late tickets A ticket manager | **4.8× fewer** | 5 → 2 | 2 October |
| Log in, add 3 contacts, list them Microsoft Playwright MCP | **3.9× fewer** | 12.6 → 4 | 2 October |
| A documentation question Context7 | **1.2× fewer** | 4 → 3 | 5 October |

The more steps a task has, the more Delta MCP saves.

[All measurements and the method](https://www.mcpdelta.com/benchmarks)

### Try Delta MCP on your own tasks

Free, two minutes to set up, and you can put everything back in one click.

[Download for Mac](https://www.mcpdelta.com/download)

Free · No account · Put everything back in one click

## Do I still need Delta MCP if I use Tool Search?

Yes, when your tasks have several steps. Tool Search keeps the menu small, but every step is still a pass. Delta MCP removes those passes: your AI writes the task as one program. Our measurements were taken with Tool Search on, so the savings come on top of it.

## Questions

### Why does MCP use so many tokens?

Two reasons. The model reads every tool definition before it can choose one, and every tool call is a separate request that re-reads the whole conversation, with each result flowing back through the model. The first is the menu tax, the second the round-trip tax.

### How many tokens do MCP tools use?

It depends on your servers. Anthropic measured five servers with 58 tools at about 55,000 tokens before the first message, and GitHub’s 35 tools alone at about 26,000. Claude Code’s `/context` command shows your own number.

### Does Tool Search fix MCP token usage?

It fixes the menu, not the round trips. Anthropic reports 85% fewer tokens for the menu, and Cursor’s equivalent cut tokens by 46.9% in runs that called MCP tools. But each tool call is still its own request over the whole conversation.

### Is code execution with MCP really cheaper?

On multi-step tasks, often. Anthropic’s example fell from 150,000 to 2,000 tokens, and Cloudflare exposed more than 2,500 API endpoints in about 1,000 tokens. Both come with a cost: you need a safe place to run the code. In our own measurements Delta MCP used 3.3–24.1× fewer tokens on multi-step tasks.

### Does Delta MCP save tokens?

Yes, on tasks with several steps: 3.3–24.1× fewer tokens in our tests, because your AI writes one program instead of one call per step.

### Do I have to change my MCP servers?

No. Delta MCP uses them exactly as they are, and every original tool stays callable by name: 152 of 152 on the largest server we tried.

### How do I check how many tokens my MCP servers use?

In Claude Code, `/context` shows what fills your context window, MCP tool definitions included. `/mcp` lists your servers and lets you switch off the ones you don’t use, and on a subscription `/usage` breaks recent usage down by individual MCP server.

## Sources

1.  [**Anthropic**, Code execution with MCP: Building more efficient agents](https://www.anthropic.com/engineering/code-execution-with-mcp) (4 Nov 2025) The two problems with direct tool calls; the 150,000 to 2,000 tokens example; the 50,000-token transcript example; the sandbox warning.
2.  [**Anthropic**, Advanced tool use](https://www.anthropic.com/engineering/advanced-tool-use) (24 Nov 2025) 58 tools on five servers at about 55,000 tokens; 85% fewer tokens with the Tool Search Tool.
3.  [**Anthropic**, Tool search tool (documentation)](https://platform.claude.com/docs/en/agents-and-tools/tool-use/tool-search-tool) Claude’s accuracy at picking a tool degrades beyond 30–50 tools.
4.  [**Claude Code**, Manage costs effectively (documentation)](https://code.claude.com/docs/en/costs) MCP tool definitions deferred by default; /context, /mcp and /usage; the full conversation sent with every request; CLIs and MCP overhead.
5.  [**Cursor**, Dynamic context discovery](https://cursor.com/blog/dynamic-context-discovery) (6 Jan 2026) 46.9% fewer tokens in runs that called an MCP tool, in Cursor’s own A/B test.
6.  [**Cloudflare**, Code Mode: give agents an entire API in 1,000 tokens](https://blog.cloudflare.com/code-mode-mcp/) (20 Feb 2026) More than 2,500 endpoints in about 1,000 tokens, against 1.17 million as a plain MCP server.
7.  [**Delta MCP**, Benchmarks](https://www.mcpdelta.com/benchmarks) Every measurement of Delta MCP on this page, with its method.

## Keep reading

-   [Claude Code usage limit reached? What your MCP servers cost What Claude Code's "You've hit your limit" means (session, weekly, Opus, spend), how /usage shows what your MCP servers cost, and what to cut first.](https://www.mcpdelta.com/guides/claude-code-usage-limit-mcp)
-   [MCP vs CLI for AI agents: where the token bill comes from MCP or CLI for your AI agent? What each costs in tokens, what the measurements show, when a CLI wins, when MCP does, and where code beats both.](https://www.mcpdelta.com/guides/mcp-vs-cli)
-   [MCP config file locations: Claude, Cursor, VS Code and more Where Claude Desktop, Claude Code, Cursor, VS Code, Windsurf, Gemini CLI and Codex keep their MCP config: file path, JSON key, remote servers and tool limits.](https://www.mcpdelta.com/guides/mcp-clients)

![](https://www.mcpdelta.com/brand/app-icon.svg)

## Try Delta MCP on your own tasks

Free, two minutes to set up, and you can put everything back in one click.

[Download for Mac](https://www.mcpdelta.com/download) Also for [Windows](https://www.mcpdelta.com/download#windows) [Linux](https://www.mcpdelta.com/download#linux)

Free · No account · Apple chip or Intel
