Short answer: MCP tool-schema bloat is the cause only when the model-facing tool surface is already large before useful work begins. If growth starts after file reads, tool results, or repeated turns, a schema gateway is solving the wrong layer.

Five causes can produce the same symptom

Total usage is an accounting result, not a root-cause analysis. Start with when the request grows.

LayerSignalWhat to measure
Standing tool schemasInput is large before the first useful tool call.Count tools and serialize request/header.tools for every request.
Tool results and historyInput grows after searches, file reads, logs, or verbose calls.Measure each retained message, tool argument, and model-facing result separately.
Repeated model stepsEach request looks plausible, but the session total is large.Count model requests per turn and separate per-request usage from cumulative usage.
Reasoning and outputInput is controlled, but completion usage remains high.Keep provider-reported output and reasoning fields separate from input.
Prompt-cache missesUncached input jumps after the tool list or schemas change.Compare cache-read and uncached input per request, then correlate with tool-surface changes.

A six-layer audit finds the first source of growth

  1. Define the number. Is it one request or the whole session? Split input, cache-read input, output, and reasoning.
  2. Measure schemas. Record tool count and serialized tool-schema JSON bytes on each request.
  3. Measure history. Attribute payload size to system text, messages, tool arguments, and rendered tool results.
  4. Count steps. Map every model request to the tool calls around it.
  5. Read provider usage. Treat local JSON size as a diagnostic proxy, not a token conversion.
  6. Run one controlled change. Hold model, tasks, server, and Harness configuration constant; change only tool exposure.

Native MCP is the baseline; progressive disclosure is a trade

DeepSeek Harness is plugin-first. Its documented architecture lets a plugin extend the tool registry without patching the agent loop. That makes both paths composable, not universally equivalent.

Keep the official native client

  • The catalog is small and stable.
  • The model should receive each exact schema immediately.
  • A direct call matters more than reducing the standing tool surface.
  • You want no discovery search before the real call.

Test progressive disclosure

  • The catalog has dozens or hundreds of long-tail tools.
  • Multiple servers create a large standing schema surface.
  • Measurement shows schemas are material before work begins.
  • You can evaluate retrieval quality and the added search step.

DeepSeek's pinned MCP client documentation says discovered tools are registered as native tools and their data-dependent schema cost is paid on every request while registered. See the official client documentation and the pinned architecture. The broader MCP search-first design is also being discussed in SEP-1576.

Read the two measurements separately

Fixed 1,000-tool component fixture

647,962 B → 1,114 B measures serialized registered tool-schema JSON bytes only. It is not a provider-token, cost, latency, stability, or task-quality result. Inspect the rc.9 benchmark.

Separate three-task DeepSeek Harness pilot

Both arms completed the same three fixed tasks: 3/3 vs 3/3. The official direct client made the selected tool call; MCP Lens first called mcp_search, then mcp_call.

Observed metricOfficial direct clientMCP Lens
Completed tasks3/33/3
MCP pathDirect selected-tool callmcp_search → mcp_call
Output tokens491794

These three cases show task-completion parity and expose the added discovery work. They do not establish a general quality, latency, cost, or stability result. Inspect the complete pilot record.

What MCP Lens changes—and what it cannot fix

MCP Lens does

  • Keep the model-facing MCP surface at mcp_search and mcp_call.
  • Return a bounded candidate set with exact schemas after search.
  • Call one explicit server and tool after another policy check.
  • Start with allowTools: [], so exposure is opt-in.

MCP Lens does not

  • Shrink tool-result text or conversation history retained by Harness.
  • Reduce model reasoning or final-answer length.
  • Fix cache misses unrelated to the tool surface.
  • Guarantee lower tokens, cost, latency, stability, or better task quality.
  • Sandbox remote MCP processes or replace endpoint security.

Direct answers

Why is DeepSeek Harness using so many tokens?

High usage can come from standing tool schemas, retained tool results and history, repeated model steps, reasoning and output, or prompt-cache misses. Split those layers before changing the MCP setup.

How can I tell whether MCP tool schemas are the cause?

Measure tool count and serialized request/header.tools JSON on each request. If that surface is large before any tool result exists, and a controlled run with fewer visible tools reduces the same layer, schema bloat is a plausible cause.

Does MCP Lens guarantee lower total token usage?

No. Schema JSON bytes are not tokens. The separate three-task pilot also shows that Lens added a search step and used more output tokens in those cases.