Short answer: MCP tool-schema bloat is the cause only when the model-facing tool surface is already large before useful work begins. If growth starts after file reads, tool results, or repeated turns, a schema gateway is solving the wrong layer.
Five causes can produce the same symptom
Total usage is an accounting result, not a root-cause analysis. Start with when the request grows.
| Layer | Signal | What to measure |
|---|---|---|
| Standing tool schemas | Input is large before the first useful tool call. | Count tools and serialize request/header.tools for every request. |
| Tool results and history | Input grows after searches, file reads, logs, or verbose calls. | Measure each retained message, tool argument, and model-facing result separately. |
| Repeated model steps | Each request looks plausible, but the session total is large. | Count model requests per turn and separate per-request usage from cumulative usage. |
| Reasoning and output | Input is controlled, but completion usage remains high. | Keep provider-reported output and reasoning fields separate from input. |
| Prompt-cache misses | Uncached input jumps after the tool list or schemas change. | Compare cache-read and uncached input per request, then correlate with tool-surface changes. |
A six-layer audit finds the first source of growth
- Define the number. Is it one request or the whole session? Split input, cache-read input, output, and reasoning.
- Measure schemas. Record tool count and serialized tool-schema JSON bytes on each request.
- Measure history. Attribute payload size to system text, messages, tool arguments, and rendered tool results.
- Count steps. Map every model request to the tool calls around it.
- Read provider usage. Treat local JSON size as a diagnostic proxy, not a token conversion.
- Run one controlled change. Hold model, tasks, server, and Harness configuration constant; change only tool exposure.
Native MCP is the baseline; progressive disclosure is a trade
DeepSeek Harness is plugin-first. Its documented architecture lets a plugin extend the tool registry without patching the agent loop. That makes both paths composable, not universally equivalent.
Keep the official native client
- The catalog is small and stable.
- The model should receive each exact schema immediately.
- A direct call matters more than reducing the standing tool surface.
- You want no discovery search before the real call.
Test progressive disclosure
- The catalog has dozens or hundreds of long-tail tools.
- Multiple servers create a large standing schema surface.
- Measurement shows schemas are material before work begins.
- You can evaluate retrieval quality and the added search step.
DeepSeek's pinned MCP client documentation says discovered tools are registered as native tools and their data-dependent schema cost is paid on every request while registered. See the official client documentation and the pinned architecture. The broader MCP search-first design is also being discussed in SEP-1576.
Read the two measurements separately
Fixed 1,000-tool component fixture
647,962 B → 1,114 B measures serialized registered tool-schema JSON bytes only. It is not a provider-token, cost, latency, stability, or task-quality result. Inspect the rc.9 benchmark.
Separate three-task DeepSeek Harness pilot
Both arms completed the same three fixed tasks: 3/3 vs 3/3. The official direct client made the selected tool call; MCP Lens first called mcp_search, then mcp_call.
| Observed metric | Official direct client | MCP Lens |
|---|---|---|
| Completed tasks | 3/3 | 3/3 |
| MCP path | Direct selected-tool call | mcp_search → mcp_call |
| Output tokens | 491 | 794 |
These three cases show task-completion parity and expose the added discovery work. They do not establish a general quality, latency, cost, or stability result. Inspect the complete pilot record.
What MCP Lens changes—and what it cannot fix
MCP Lens does
- Keep the model-facing MCP surface at
mcp_searchandmcp_call. - Return a bounded candidate set with exact schemas after search.
- Call one explicit server and tool after another policy check.
- Start with
allowTools: [], so exposure is opt-in.
MCP Lens does not
- Shrink tool-result text or conversation history retained by Harness.
- Reduce model reasoning or final-answer length.
- Fix cache misses unrelated to the tool surface.
- Guarantee lower tokens, cost, latency, stability, or better task quality.
- Sandbox remote MCP processes or replace endpoint security.
Direct answers
Why is DeepSeek Harness using so many tokens?
High usage can come from standing tool schemas, retained tool results and history, repeated model steps, reasoning and output, or prompt-cache misses. Split those layers before changing the MCP setup.
How can I tell whether MCP tool schemas are the cause?
Measure tool count and serialized request/header.tools JSON on each request. If that surface is large before any tool result exists, and a controlled run with fewer visible tools reduces the same layer, schema bloat is a plausible cause.
Does MCP Lens guarantee lower total token usage?
No. Schema JSON bytes are not tokens. The separate three-task pilot also shows that Lens added a search step and used more output tokens in those cases.