Same result, up to 24.1× fewer tokens.
With Delta MCP, multi-step tasks use far fewer tokens, cost less and finish as fast or faster, with the same end result, checked in the service itself. 8 scenarios, each run with and without Delta MCP, in Claude Code with Sonnet 5.5.
fewer billed tokens on multi-step tasks
lower cost on multi-step tasks, even in the strictest cache case
faster on multi-step tasks, and no scenario is slower
tools still declared and callable through Delta MCP
Every scenario, every number
From a release fix that takes 26 model calls without Delta MCP, down to a single question. Averages of 2 to 3 runs per variant, with the same end result in every scenario. “Weighted cost” is the cost Claude Code reports per model, in input-token equivalents, in the strictest cache case. “As code” is the default for every server; “as text” uses a Cerberus adapter.
| Task | Server | Tokens | Weighted cost | Model calls | Time | Measured |
|---|---|---|---|---|---|---|
| Fix a release, as text | Cerberus, 152 tools | 24.1× fewer | 8.1× fewer | 26 → 4.5 | 1.2× faster | 2 Oct |
| Add tests, as text | Cerberus, 152 tools | 15.8× fewer | 5.8× fewer | 16.6 → 4 | 2.0× faster | 2 Oct |
| Fix a release, as code | Cerberus, 152 tools | 13.4× fewer | 6.1× fewer | 26 → 6.5 | as fast | 5 Oct* |
| Clean up 20 late tickets | A ticket manager | 4.8× fewer | 4.0× fewer | 5 → 2 | 4.4× faster | 2 Oct |
| Log in, add 3 contacts, list them | Microsoft Playwright MCP | 3.9× fewer | 1.8× fewer | 12.6 → 4 | 2.1× faster | 2 Oct |
| Make 6 changes to a task list | taskqueue-mcp | 3.3× fewer | 1.5× fewer | 5.6 → 2 | 2.1× faster | 2 Oct |
| A documentation question | Context7 | 1.2× fewer | 1.11× more | 4 → 3 | as fast | 5 Oct |
| A one-step question | Playwright, books.toscrape.com | 1.25× more | 2.5× more | 4 → 3 | 1.1× faster | 5 Oct |
* Delta MCP measured on 5 October, direct calls on 2 October.
It holds up
Holds on tasks it never saw
Missions written after the tuning, and never used to tune it, used 3.7 to 4.1 times fewer tokens.
The same quality, checked run by run
26 of 27 runs were flawless in the final round, and the 27th had nothing false in it. All 103 runs were read one by one: exact texts, no field touched by mistake, a right answer to the user. The audit also caught two defects the automatic checks had missed, now fixed.
How we measure
One method for every number on this page, each one traceable to its raw runs.
- Billed tokens, thinking included, as the provider counts them.
- A run counts only if the requested model actually served it.
- Model calls are distinct messages from the model, not tool calls.
- Success is the end state checked in the service, and the answer to the user.
- Cost in the strictest cache case; every gain rounded down.
- Every published figure has an entry in a claims register, recomputed from the raw runs.
Next: the biggest MCP servers
Delta MCP will be put to the test on the biggest MCP servers, and the results will be published when it is released. Be the first to see them.
Keep reading
-
How to reduce MCP token usage, and why MCP burns tokens
How to reduce MCP token usage: measure it with /context, see the two costs (tool menu and round trips), and compare every fix, Tool Search included.
-
MCP vs CLI for AI agents: where the token bill comes from
MCP or CLI for your AI agent? What each costs in tokens, what the measurements show, when a CLI wins, when MCP does, and where code beats both.
-
Claude Code usage limit reached? What your MCP servers cost
What Claude Code's "You've hit your limit" means (session, weekly, Opus, spend), how /usage shows what your MCP servers cost, and what to cut first.
Measure it on your own tasks
Free, on your computer. Delta MCP keeps count of the tokens it saves you.
Free · No account · Apple chip or Intel
Free · No account
Free · No account
For Mac, Windows and Linux. No account needed.