The short answer
A CLI usually costs fewer tokens than an MCP server: it adds no tool menu, and the model already knows common commands. MCP wins where no CLI exists, where you need sign-in, or where your app has no shell. In both, every step is a request. Chaining the steps in code removes that cost.
In this guide
What does each one cost in tokens?
Three things drive the bill: the menu of tools, the number of steps, and what comes back from each step. MCP servers and CLIs differ on all three.
| Cost | MCP server | CLI |
|---|---|---|
| The menu | Every tool has a definition the model must read. Claude Code defers them by default: only tool names enter the context until a tool is used. | None. The model already knows common commands, or reads --help once. |
| Each step | One tool call is one request over the whole conversation. | One command is one request too, but a command can chain several steps with pipes and &&. |
| What comes back | The whole result enters the conversation. Claude Code warns above 10,000 tokens and caps a result at 25,000 by default. | You can trim it before the model sees it (head, jq), or save it to a file, as Playwright CLI does. |
Chaining and trimming come from the shell, not from any one CLI: it is a small program that runs several steps in one request. Code mode, below, gives MCP the same ability.
What do the measurements say?
Three sources, from three angles.
- AnthropicClaude Code’s cost guide says CLI tools such as
gh,aws,gcloudandsentry-cliare still more context-efficient than MCP servers, because they add no per-tool listing. - MicrosoftThe Playwright CLI README says CLI invocations are more token-efficient than its MCP server: they avoid loading large tool schemas and verbose accessibility trees into the model. It keeps MCP for exploratory automation, self-healing tests and long-running workflows, where continuous browser context matters more than token cost.
- ScalekitA benchmark from Scalekit, March 2026: five read-only GitHub tasks, 75 runs, Claude Sonnet 4. The
ghCLI used 1,365 to 9,386 tokens per task and GitHub’s MCP server 32,279 to 82,835, which is 4 to 32 times more. The MCP server also failed 7 of 25 runs with timeouts; the CLI succeeded every time.
That benchmark loaded all 43 of GitHub’s tool definitions into every conversation. Claude Code now defers them by default, which removes the menu part of that gap there, but not the per-step part, and not in apps that still load every definition.
When does a CLI win, and when does MCP?
A CLI wins when…
- The tool is a well-known one, such as gh, aws or gcloud.
- The output is long and you need only part of it: trim it before the model reads it.
- You work in a coding agent that has a shell and a filesystem.
- Several steps fit in one command.
MCP wins when…
- The service has no CLI, or only a remote server with its own sign-in.
- Your AI app has no shell for the model to use, as in a chat app.
- You want typed inputs and a list of what each tool does, instead of flags the model recalls from training.
- Several people use the agent and each needs their own authorization: Scalekit’s main point.
Is MCP dead?
No. The protocol still does what a CLI cannot: reach remote services with sign-in, work in apps that have no shell, and describe each tool in a typed way. Microsoft’s own README keeps MCP for some workflows.
What is wearing out is a pattern: one tool call per step, over the whole conversation. A CLI avoids part of it by chaining commands. Anthropic and Cloudflare describe the general fix, writing code that calls the tools, and measure it: Anthropic’s example fell from 150,000 to 2,000 tokens, and Cloudflare exposed more than 2,500 endpoints in about 1,000 tokens.
Delta MCP is that fix, ready-made, for the MCP servers you already use: your AI writes each task as one program. We measured it against calling those servers directly, on Claude Code with Tool Search on:
Sonnet 5.5, two to three runs per variant, the end result checked in each service.
| Task | Tokens | Model calls | Measured |
|---|---|---|---|
| Fix a release, as codeCerberus, 152 tools | 13.4× fewer | 26 → 6.5 | 5 October |
| Clean up 20 late ticketsA ticket manager | 4.8× fewer | 5 → 2 | 2 October |
| Log in, add 3 contacts, list themMicrosoft Playwright MCP | 3.9× fewer | 12.6 → 4 | 2 October |
All measurements and the method
Which should you use?
Start from your situation, not from the protocol.
| Your situation | Use | Why |
|---|---|---|
| The tool is well known and has a CLI: gh, aws, gcloud. | A CLI | No menu, and you can trim the output. Claude Code’s cost guide recommends it. |
| The service has no CLI, or you reach it through a remote server. | MCP | It is the way in. Keep the servers you use and disconnect the rest. |
| Your AI app has no shell. | MCP | A CLI needs somewhere to run. |
| A task crosses several steps or several services. | Code, not one call per step | Delta MCP does it for your MCP servers. A script does it for CLIs. |
Questions
Is MCP or CLI better for AI agents?
Neither always. A CLI is usually cheaper where a well-known one exists, because it adds no tool menu and its output can be trimmed. MCP is the way in where no CLI exists, where sign-in is needed, or where the app has no shell. Both pay a request for every step.
Is a CLI cheaper than an MCP server in tokens?
Usually, per task. Scalekit measured the gh CLI at 1,365 to 9,386 tokens per task against 32,279 to 82,835 for GitHub’s MCP server, with all 43 tool definitions loaded. Claude Code now defers those definitions, which narrows the gap.
Is MCP dead?
No. It still covers remote services, sign-in and apps without a shell, and Microsoft’s own Playwright CLI README keeps MCP for exploratory and long-running workflows. What costs tokens is one call per step, which code can replace.
Should I replace my MCP servers with skills and CLIs?
Only where a good CLI exists. Skills help too: Claude Code loads only their descriptions at the start and the full text when one is invoked, and a skill can include code to run. Anthropic says skills complement MCP servers rather than replace them.
Why is there a Playwright CLI next to the Playwright MCP server?
Its README says CLI invocations avoid loading large tool schemas and verbose accessibility trees into the model, and that CLI plus skills suits coding agents that balance browser automation with large codebases in limited context windows. MCP stays for exploratory automation, self-healing tests and long-running workflows.
Does Delta MCP replace a CLI?
No. A CLI is the right tool where a good one exists. Delta MCP is for the MCP servers you already have, on tasks with several steps: 3.3–24.1× fewer tokens in our tests against calling them directly.
Sources
- Claude Code, Manage costs effectively (documentation) CLI tools are more context-efficient than MCP servers; MCP tool definitions are deferred by default.
- Claude Code, Connect Claude Code to tools via MCP (documentation) The 10,000-token warning and the 25,000-token default cap on a tool result.
- Microsoft, Playwright CLI (README on GitHub) CLI against MCP on token efficiency, and where MCP remains the better fit.
- Scalekit, MCP vs CLI: Benchmarking AI Agent Cost & Reliability (11 Mar 2026) The GitHub benchmark: tokens per task, reliability and the 43 tool definitions.
- Anthropic, Code execution with MCP: Building more efficient agents (4 Nov 2025) The 150,000 to 2,000 tokens example; code that calls tools instead of one call per step.
- Cloudflare, Code Mode: give agents an entire API in 1,000 tokens (20 Feb 2026) More than 2,500 endpoints in about 1,000 tokens, against 1.17 million as a plain MCP server.
- Anthropic, Equipping agents for the real world with Agent Skills (16 Oct 2025) Skills load their name and description first and the rest on demand; they can include code; they complement MCP servers.
- Claude Code, Extend Claude with skills (documentation) Skill descriptions load at the start; the full content only when a skill is invoked.
- Delta MCP, Benchmarks Every measurement of Delta MCP on this page, with its method.
Keep reading
- How to reduce MCP token usage, and why MCP burns tokensHow to reduce MCP token usage: measure it with /context, see the two costs (tool menu and round trips), and compare every fix, Tool Search included.
- Claude Code usage limit reached? What your MCP servers costWhat Claude Code's "You've hit your limit" means (session, weekly, Opus, spend), how /usage shows what your MCP servers cost, and what to cut first.
- MCP config file locations: Claude, Cursor, VS Code and moreWhere Claude Desktop, Claude Code, Cursor, VS Code, Windsurf, Gemini CLI and Codex keep their MCP config: file path, JSON key, remote servers and tool limits.