Skip to content
robinSenior Software Engineer
All articlesReducing MCP Token Usage by 85%

Reducing MCP Token Usage by 85%

3 Apr 2026 · 3 min read · 519 words

A direct GitHub MCP connection can put 24,473 schema tokens in context before the model makes a call. A 200-item API response can add 312,877 bytes. I built tldr to reduce both costs.

Two costs, one wrapper

MCP context has two costs. The first is schema injection. The second is response volume.

Connect a coding harness directly to several upstream MCP servers and the model sees every tool definition before it has a task. Then, when a tool runs, its full JSON response often goes straight back into context. The costs compound over a long session.

tldr is a local gateway between the harness and those servers. The harness connects to one server, tldr serve, and asks it to discover or call upstream tools. The wrapper exposes five small operations: search_tools, execute_plan, call_raw, inspect_tool, and get_result.

Loading diagram...

Defer tool discovery

The wrapper compiles each upstream tool into a smaller Capability record with its name, summary, tags, risk level, input shape, and output shape. The model gets candidates first, then requests a full schema only when it needs one.

The GitHub MCP server is a useful measurement. Its 41 tools serialise to about 24,473 schema tokens. The compiled capability index is about 3,482 tokens, an 85.77% reduction before the first tool call.

Loading diagram...

Those numbers use the same rough estimator: one token per four bytes of serialised JSON. They are not universal tokenizer counts. They are a consistent comparison between two representations.

Shield responses and keep the raw result

Schema compression only solves the first half. A 300 KB response can erase the saving on the next turn.

tldr stores each raw result locally, then returns a smaller view. Arrays are capped at 50 elements, long strings at 8,192 characters, and valid JSON is summarised before the system falls back to byte truncation. The untouched result remains available under a reference such as p1:s1.

A representative 200-item GitHub-style issues response measured 312,877 bytes before shielding and 85,620 bytes after it. That is a 72.63% reduction. The main saving came from trimming the array to 50 summarised items.

Loading diagram...

Tool compression saves prompt budget before execution. Response shielding saves it after execution. You need both.

Fetch the exact fragment

The result reference keeps shielding from being destructive. get_result supports pagination with offset and limit, field projection with fields, nested navigation with path, and search with pattern, before, and after.

The store materialises searchable text into a cache file on the first search, then reuses it. Search results and their cache files expire together. Results live for one minute by default, and in-memory storage is capped at 128 MB.

That is the useful interaction: return a cheap summary first, then retrieve only the lines the model needs.

Loading diagram...

Measure your own setup

The measured savings depend on the tools and responses in your stack. Measure schema tokens before the first call, response bytes after execution, and how often the agent needs to retrieve an omitted fragment.

Install the wrapper with:

curl -sSfL https://raw.githubusercontent.com/robinojw/tldr/main/install.sh | sh

Then wrap a real MCP server and point your harness at tldr serve. A gateway earns its place only when it reduces total context cost without making the agent less capable.