Why did loading all my MCP tools blow the context window
Every MCP tool schema you attach is prompt tokens. Loading the whole catalog fills the window before the task starts—send a job-sized subset or a router.
William Spurlock Founder — Spurlock Studios 32 MIN
Loading every MCP tool blew the context window because the host stuffed each tool’s name, description, and JSON Schema into the model request. That tax is billed before the user sentence. You did not run forty tools. You advertised forty tools. The model had to read the zoo in order to pick one cage.
This is not the planning-loop tax of later turns. That bill is turns × a growing transcript. This bill is the prefix: schemas sitting in context on turn one. It belongs under the Agentic Systems Operating Manual. If the job is a known lookup with one write, you may not want an agent at all — that brake lives in when not to build an agent.
The short answer
- Schemas are context. OpenAI’s function-calling guide is blunt: callable definitions count against the context limit and are billed as input tokens.
- MCP discovery ≠ the prompt. As of the 2026-07-28 tools spec,
tools/listtells the client what exists. The host still chooses what the model sees this turn. - Load a subset or a router. Job-typed allowlist first. Vendor
defer_loading/ tool search, or asearch_toolsmeta-tool, when the measured prefix still crowds the task. - Do not quote a fake constant. Token cost depends on description length, nested JSON Schema, and whether the host forwards
outputSchemaand other metadata. Count your serialized payload. - Results are tax two. A fat ticket dump on turn one will blow the window even after you shrink the catalog.
Why did a one-line task still blow the window?
Because the user message was never the bulk of the request. The bulk was the tool catalog the host attached so the model could “have access.” Access is not free. The model has to read every schema you sent in order to decide which one to call.
Anthropic’s code execution with MCP (Nov 4, 2025) names the two overload patterns: tool definitions loaded upfront, and intermediate results passed through the model. Their write-up describes agents connected to hundreds or thousands of tools across dozens of servers. Treat that as vendor framing of the failure, not a census of your stack.
| What you thought happened | What the request actually contained | Who pays |
|---|---|---|
| “I asked one question” | System prompt + every attached tool schema + the question | Input tokens, turn 1 |
| “I have not called GitHub yet” | GitHub’s full inputSchema set is already in the prompt | Same |
| “MCP is just a connector” | The host copied tools/list into the provider tools array | Host policy, not the spec |
| “The window is 200k, we are fine” | Prefix of tens of thousands of tokens leaves little room for the task, history, and results | You, when the completion truncates or the model picks the wrong twin tool |
A one-line task with an 80-tool prefix is a small question taped to a large dictionary. The dictionary is the bill.
What actually entered the model request?
Pin the layers. People say “MCP used all my context” when they mean “my host forwarded every listed tool into the provider request.” Those are different objects.
| Layer | What it holds | Does it occupy the model window? |
|---|---|---|
| MCP server process | Handlers, secrets, real side effects | No |
tools/list RPC | Name, description, inputSchema, optional outputSchema, icons, annotations | Only if the host copies it onward |
| Host allowlist / router | The subset this job may see | Indirectly: it decides the next row |
Provider tools / input_schema array | The schemas the model is allowed to call now | Yes. This is the prefix tax. |
| Tool result on the next turn | JSON / text the handler returned | Yes. This is tax two. |
Procedure for a blown window:
- Export the raw model request (not the MCP inspector view).
- Isolate the
tools/ functions array from messages. - Serialize that array exactly as the SDK sent it.
- Count tokens with the same tokenizer the provider bills, on that payload alone.
- Compare that number to the user message. If the catalog dwarfs the question, you found the leak.
If you cannot export the request, you are guessing. Guessing is how “always 500 tokens per tool” folklore gets invented. Kill the folklore. Count the bytes you actually sent.
Native function calling pays the same prefix tax. MCP did not invent “schemas in the prompt.” MCP made it easy to attach twenty servers you would never have typed by hand. The protocol is the delivery truck. The host is who loaded the truck.
| How the window “blows” | What operators see | What it is |
|---|---|---|
| Hard overflow | Provider 400 / context-length error on turn 1 | Prefix + messages exceeded the model limit |
| Silent squeeze | Request succeeds; task, history, or results get truncated by the host | You “fixed” the error by dropping the work |
| Quality collapse | Completes, wrong tool, ignores the question | Catalog crowding; twins in the prompt |
| Cache thrash | Stable system prompt, but every turn misses the cache | Tool array order or membership changed |
“Blew the window” is any of those four. Only the first looks like a red error. The other three still waste money and miss the ticket.
Is tools/list the same as the catalog the model sees?
No. Mixing those two is the usual architecture error.
The 2026-07-28 tools spec (confirm this line on the current public docs before you promise a client a feature) says servers that declare tools must answer tools/list. The set may be empty, may change over time, must not vary per-connection as a side effect of other requests, and may vary by the authorization on the request. Listing supports pagination and caching. Servers should return a deterministic order so clients can cache and so prompt-cache hit rates improve when tools are included in model context.
That last clause is the tell. The spec knows tools often land in the model prompt. It does not require the host to include every listed tool.
| Job | Consumer | Healthy behavior |
|---|---|---|
tools/list (paged) | Host / gateway | Learn what the server could run |
| Auth-scoped list | Same | Hide tools the token cannot use |
| Per-run allowlist | Host | Choose verbs for this job |
| Provider tool array | The model | Only the allowlist, or a router plus a handful of always-on tools |
Pagination does not save the window by itself. A client that drains every nextCursor and then forwards the union to the model has a tidy RPC and a ruined prompt. Cursor opacity, server-chosen page size, missing limit= in the protocol — those are wire facts. They are not a catalog policy.
If your “fix” was “the server paginates tools/list now,” you fixed the JSON-RPC payload. You have not fixed the model request until the host stops stuffing the union.
listChanged (current tools spec) means the server may notify that the list moved. Useful for the host cache. Dangerous if the host reacts by re-dumping a new union into the next completion and breaking the prompt cache. Treat a list-changed notification as “refresh the allowlist intersection,” not “send everything again.”
Hosts that also dump resources and prompts into the model are running a second zoo. MCP paginates resources/list and prompts/list for the same reason it paginates tools. The window still dies if the host inlines resource bodies or every prompt template on turn one.
| MCP surface | List method (current spec) | Model-window rule |
|---|---|---|
| Tools | tools/list | Attach the job subset, not the union |
| Resources | resources/list | Fetch by id when the job needs the blob |
| Prompts | prompts/list | Pick the template for this job; do not paste the library |
| Templates | resources/templates/list | Same as resources: discover ≠ inline |
Confirm those method names on the live spec if you are pinning an older line. Do not assume 2025 SSE-era client code still matches.
Why are schemas expensive when nobody called a tool?
Because the model cannot call what it cannot see, and seeing is reading. A tool the handler never ran still occupied context on every turn you attached it.
Cost drivers that show up in real catalogs (ranges, not constants):
| Driver | Why tokens climb | What to do |
|---|---|---|
Long description | Prose is tokens. Essays in the description field are prompt bloat. | Write “when to pick this tool,” not the handbook. |
| Nested JSON Schema | Every properties key, enum, and $ref expansion is text. | Keep arguments the handler actually reads. |
| Duplicate twins | jira_create_issue vs jira_create_ticket vs create_issue | One verb per side effect. |
outputSchema forwarded | Optional on current MCP tools spec; hosts that copy it add more prompt. | Keep output schema for validation in the host if you need it; do not assume the model must eat it. |
| Examples inside the schema | Few-shot in JSON is still tokens every turn. | Put examples in evals, or in a deferred doc the router can fetch. |
| Icons / titles / extra metadata | Harmless as URLs; painful if a host inlines blobs. | Do not inline image bytes into the tool array. |
Same tool, two description styles. Count these strings with your tokenizer. Do not reuse whatever number you get as a constant for GitHub’s server.
Fat (handbook in the schema — avoid this):
{
"name": "get_issue",
"description": "This tool retrieves a Jira issue from the configured cloud site. Jira is Atlassian’s issue tracker. Use this whenever the user mentions a ticket, incident, bug, story, or epic. The issue key looks like PROJ-123. Never guess. The response includes the full rendered HTML of the description, all comments, all attachments metadata, watchers, sprint, and custom fields. If you are unsure whether to use this or search_issues or create_issue or transition_issue, consider the user’s intent carefully and prefer this tool for lookups.",
"inputSchema": {
"type": "object",
"properties": {
"issue_key": { "type": "string", "description": "The issue key." },
"expand": { "type": "string", "description": "Optional expand query. You can pass many comma-separated values." },
"include_html": { "type": "boolean" },
"include_comments": { "type": "boolean" },
"include_attachments": { "type": "boolean" }
}
}
}
Thin (when to pick it — prefer this):
{
"name": "get_issue",
"description": "Fetch one Jira issue by key (e.g. INC-1842). Use for status lookups. Do not use to create or transition issues.",
"inputSchema": {
"type": "object",
"additionalProperties": false,
"required": ["issue_key"],
"properties": {
"issue_key": {
"type": "string",
"description": "Existing Jira key from the user or a prior search. Never invent a key."
}
}
}
}
The fat version is still one tool. It still belongs in the prefix on every turn you attach it. Multiply that style across forty verbs and you will not need a vendor anecdote to know why the window died.
OpenAI’s function-calling page says that if you hit token limits, limit the number of functions loaded up front, shorten descriptions, or use tool search so deferred tools load only when needed. That is the vendor admitting the zoo is expensive. It is not a license to invent “tool X always costs N.”
Anthropic’s advanced tool use post (Nov 24, 2025) published a worked example of one five-server mix: GitHub 35 tools (~26K tokens), Slack 11 (~21K), Sentry 5 (~3K), Grafana 5 (~3K), Splunk 2 (~2K) — about 58 tools and ~55K tokens before the conversation starts, in their catalog. Jira in that post is ~17K alone. They also report seeing tool definitions consume 134K tokens before optimization. Those are measurements of specific servers on a specific date. Recite them as Anthropic’s example. Do not paste 55K into a customer deck as if it were your GitHub server.
How do desktop and IDE hosts dump the zoo?
Most people meet this bug in a desktop or IDE host, not in a custom worker. You tick twelve MCP servers in settings because they might be useful. The host connects all of them, lists every tool, and — unless it has a defer path — attaches the union to the next completion.
Custom workers repeat the same mistake with more confidence: await Promise.all(servers.map(listTools)), then chat.completions.create({ tools: everything }).
| Host shape | Typical dump | What “subset” means here |
|---|---|---|
| Desktop / IDE MCP panel | Every enabled server, every listed tool, every session | Disable servers you are not using this hour; prefer hosts that defer |
| Overnight agent worker | Every server in the deploy manifest, every job type | Filter by job type in code before the model call |
| Multi-tenant gateway | Every tool the caller’s token can see | Scope the token and the prompt; auth-scoped tools/list is not enough if you still forward the list |
| Eval harness | Golden-set tools plus leftovers from a local config | Pin the catalog in the eval fixture |
Checklist for the host you actually run:
- I can name which servers are connected in production, not just in my laptop config
- I can name which of those servers this job type is allowed to see
- I can show the
toolsarray on a trace for that job type - I have a way to disable a noisy server without a full redeploy
- I have not copied my desktop MCP config into the worker “so prod matches my IDE”
Desktop convenience is not a production catalog. If the worker’s first request includes Slack, Sentry, Grafana, and a filesystem server so it can “be helpful,” you paid the zoo tax on a job that needed one ticket tool.
Failure mode: forty servers, one ticket lookup
What breaks: A support agent is asked “What is the status of INC-1842?” The host has GitHub, Jira, Slack, Confluence, Sentry, a SQL server, a browser server, and a filesystem server connected. The model spends the prefix reading create-issue, close-issue, list-channels, run-query, and write-file. It then either (a) overflows, (b) picks a lookalike write tool, or (c) answers from a stale description because the right get_issue schema was a needle in the stack.
What it costs: Input tokens on every turn, worse tool selection, and a smaller remainder for the actual ticket body. If the first successful call returns the full issue JSON including comments, attachments, and HTML, tax two finishes the job the schemas started.
What you do instead:
- Job type
support.status_lookup→ allowlist{jira.get_issue}(read). - If the run must comment, that is a second job type with
{jira.get_issue, jira.add_comment}and a policy gate. - Writes still use runtime idempotency keys — see idempotent agent tool writes. A smaller catalog does not make a double-post safe.
- Cap the get_issue result: identifiers, status, assignee, last comment. Not the attachment bytes.
| Symptom | Likely layer | First fix |
|---|---|---|
| Request 400 / context-length error on turn 1 | Prefix schemas | Drop servers; allowlist |
Completes but calls create_issue for a status question | Twin tools in the prompt | Kill the write verb for this job type |
| Turn 1 fine, turn 2 overflows | Tool result | Truncate / paginate / digest |
| Cache miss every turn | Catalog reshuffled | Deterministic tool order (spec should); stable allowlist |
| “It worked in the IDE” | Desktop had three servers; worker has twenty | Pin catalogs per environment |
This is the failure we keep seeing after 20,000+ hours on agentic systems: the demo used two tools. Production inherited the founder’s MCP panel.
How do I measure the tax without inventing a constant?
You measure the serialized catalog the host sends, plus the result payloads the loop appends. You do not average a blog-post token number across vendors.
| Metric | How to take it | Pass heuristic (yours to set) |
|---|---|---|
| Prefix tool tokens | Tokenizer on the tools array only | Small vs the model’s context; leave room for the task |
| User+system tokens | Same request, messages only | Should dominate the question, not lose to the zoo |
| Result tokens / turn | Each tool payload as sent back to the model | Cap; Anthropic notes Claude Code defaults a 25k-token tool-response cap — that is their product default, not a spec limit |
| Turns until overflow | Trace of a golden job | Job completes with remainder |
| Wrong-tool rate | Eval on lookalike names | Drops after you remove twins |
Numbered measurement procedure:
- Pick one golden job (the ticket question, not “be a general assistant”).
- Run it with the current zoo. Export the request.
- Count prefix tool tokens. Record wrong-tool and overflow.
- Cut to the allowlist you believe the job needs. Repeat the same golden job.
- If prefix is still huge, shorten descriptions and split servers. Count again.
- Only then consider a router /
defer_loading. Routers add a search turn; you should know the static allowlist was not enough.
Checklist so the number stays honest:
- Same model pin for before/after
- Same tokenizer as billing, not a guess from a third-party widget
- Catalog taken from the request, not from
tools/listin the inspector - Result caps measured separately from schema tokens
- You wrote the allowlist down so next week’s demo cannot silently reattach Slack
If someone quotes “MCP tools always cost N tokens,” ask for the serialized payload. No payload, no number.
Worksheet to paste into the design doc (fill with your traces):
| Field | Zoo run | Allowlist run | Notes |
|---|---|---|---|
| Date / model pin | Same pin both columns | ||
| Servers connected | Names, not “all” | ||
| Tools in provider array | Count of schemas | ||
| Prefix tool tokens | Tokenizer = billed | ||
| User+system tokens | |||
| Result tokens turn 1 | After the first call | ||
| Wrong-tool? (Y/N + name) | |||
| Overflow? (Y/N + turn) | |||
| Job pass/fail | Same golden prompt |
Empty cells mean you did not measure. Empty cells are not “probably fine.”
What does a job-sized subset look like?
A job-sized subset is the verbs this job is allowed to use, not the company’s platform catalog. Three to eight tools is a common production shape for a single job type. That is a design default, not a law. A lookup with one read tool is better than eight. A run that truly needs twelve related reads should still not include an unrelated write.
| Job type | Servers connected | Tools the model sees | Tools the server still owns |
|---|---|---|---|
support.status_lookup | Jira | get_issue | create, transition, attach, search |
support.comment | Jira | get_issue, add_comment | create, delete |
release.notes | GitHub | list_merged_prs | every other GitHub verb |
ops.page | PagerDuty | get_incident, ack_incident | create, delete, list-everything |
Allowlist procedure:
- Name the job type in the runner (string, not a vibe).
- Map job type → server set (usually one).
- Map job type → tool name set (usually a handful).
- Intersect with what
tools/listreturned so a renamed server fails closed. - Pass that intersection to the model. Log the names. Fail the run if the intersection is empty.
- Every production job type has a written allowlist
- Write verbs are absent from read jobs
- Omnibus
do_anything/call_apitools are not on the list - The allowlist is in the worker, not in a prompt paragraph the model can ignore
- Changing the allowlist is a reviewed config change
The server can keep a large implementation surface. The model should not.
When do I want a router instead of a static allowlist?
When one worker honestly serves many job types and a static list per type is either huge or constantly wrong. A router is a search step: the model (or the host) retrieves a handful of schemas, then calls them. It is extra latency. Pay it when you have measured that the static prefix is the problem.
| Pattern | What the model sees first | Who searches | When it fits |
|---|---|---|---|
| Static allowlist | 3–8 full schemas | You, at job-type routing time | Default. Most pilots. |
| Host pre-filter | Schemas for this tenant × job | Your gateway, no extra model call | Multi-tenant, known verbs |
OpenAI tool_search + defer_loading | Server/namespace name+description; schemas later | Provider or your client-executed search | Large declared inventory; confirm current model support on OpenAI’s tool-search guide |
Anthropic Tool Search Tool + defer_loading | Search tool + any non-deferred tools | Claude searches, then full schemas load | MCP-heavy hosts; confirm the current Claude tool-use docs, not only the Nov 2025 announcement |
| Code / filesystem progressive disclosure | Directory of servers, read schemas on demand | The model reads files or a search_tools tool | You already run a sandbox; Anthropic’s MCP code-execution post reports a worked example dropping 150,000 → 2,000 tokens in that setup — their numbers, not yours |
OpenAI’s tool-search guide (as of the public page cited here) recommends grouping deferred functions into namespaces or MCP servers with clear high-level descriptions, and treating “fewer than 10 functions per namespace” as a best-practice for their search quality. Confirm that guidance on the live page. Do not treat “10” as an MCP spec limit. There is none.
Anthropic’s same Nov 2025 post suggests considering Tool Search when tool definitions exceed ~10K tokens, when selection accuracy is already bad, when multiple MCP servers are in play, or when 10+ tools are available — and says it is less useful when the library is small or every tool is used every session. Again: their product heuristic. Your measurement from the previous section outranks it.
Router checklist:
- Static allowlist already lost on a measured prefix or wrong-tool rate
- Search corpus is names and descriptions that a retrieval step can rank
- Write tools cannot be retrieved into a read job without a policy gate
- You budgeted one extra turn for search
- Prompt cache: deferred tools should not scramble a stable prefix (OpenAI documents injecting discovered tools at the end of the window for that reason — verify current behavior)
If you cannot name the job types, a router will retrieve a slightly different zoo every turn. That is not a fix. That is a slot machine with schemas.
Hosted vs client-executed search (vendor names differ; the split is the same):
- Hosted: You declare the inventory (functions, namespaces, or MCP servers) with
defer_loadingwhere the vendor supports it, add the search tool, and let the API choose what to hydrate. Confirm the current request shape on OpenAI’s tool-search page or Anthropic’s current tool-use docs before you copy a year-old gist. - Client-executed: The model emits a search call. Your gateway queries an index you own (job-type allowlist, BM25, embeddings — you pick). You return the matching full schemas. Then the model calls a real tool.
- Policy after retrieval: A retrieved
delete_repostill must fail the gate on a read job. Search is not authorization. - Log both hops: search query, names returned, names actually called. If search returns twelve write twins, your index descriptions are lying.
A search_tools meta-tool you write yourself is valid. Keep it small: query in, a handful of {name, description, inputSchema} out, hard cap on hits. If it returns the whole catalog “just in case,” you built a router that dumps the zoo on purpose.
Results are a second window tax
Shrinking the catalog and then returning a 10,000-row sheet into the next message is how teams “fix MCP” and still overflow. Definitions are tax one. Results are tax two. Anthropic’s MCP code-execution post uses a 2-hour sales transcript as an example of ~50,000 additional tokens when the full document flows through the model twice (read, then write). That is an illustration of a fat payload, not a constant for every transcript.
| Result shape | Window effect | Prefer |
|---|---|---|
| Full document / sheet / log | Can exceed the remainder after schemas | Filter in the tool; return a digest + id |
Unbounded list_* | Page 1 is 20 rows; page “all” is the archive | Default limit; cursor; refuse limit > cap |
| Nested objects with HTML | Tags are tokens | Plain fields the next tool needs |
| Error tracebacks | Stack frames crowd the retry | One-line, actionable error the model can correct |
| Secrets in results | Transcript + logs now hold the key | Never return secrets; the model will see them |
Procedure for every list-shaped MCP tool:
- Required
limitwith a small default (you pick; Anthropic’s agent-tools post argues for pagination, range, filter, and truncation with sensible defaults). - Opaque cursor; never invent one on the client.
- Explicit truncation flag so the model knows to ask for the next page.
- Hard cap in the server, not only in the prompt.
- For binaries and long docs, return a handle the next tool can fetch — do not paste the file into the assistant message.
The MCP pagination spec covers tools/list, resources/list, prompts/list, and templates. It does not paginate tools/call results for you. Result envelopes are your server’s job. If a list tool returns the whole table, that is a handler bug, not an MCP omission.
Writes that survive this filter still need idempotent retries. Lookup-plus-comment is two side-effect classes. Keys stay in the runtime.
When is a workflow enough instead of an 80-tool agent?
When you can name the path. “Get INC-1842, format status, post to Slack” is a workflow with two HTTP nodes. Putting those two calls behind MCP so an agent can also create issues, delete files, and run SQL is how a status bot inherits a company-wide blast radius and a blown window.
Read when not to build an agent before you keep the loop. The catalog tax is one more reason that page exists: agents pay for optionality. Workflows do not attach 80 unused schemas.
| Signal | Workflow | Agent with a tiny catalog | Agent with a router |
|---|---|---|---|
| Path known, IDs in the trigger | Yes | No | No |
| Path varies, pass/fail writable, verbs ≤ ~8 | Maybe a hybrid | Yes | Not yet |
| Many job types, one worker, measured prefix pain | Split workers | Split job types first | Then router |
| “Might need every SaaS someday” | No | No | No |
Decision list:
- If the trigger already contains the issue id, do not search the universe.
- If a human would run the same three clicks every time, encode the clicks.
- If the model must choose among a few reads, give it those reads — not the write twins.
- If you cannot write pass/fail, you are not ready for a catalog discussion. You are not ready for an agent.
Optionality you do not use is still billed. That is the whole post.
How should I split servers and shrink descriptions?
Split by blast radius and job, not by vendor logo. One “company MCP server” with 90 tools is a gift to context bloat and to confused-deputy writes. Several small servers let a host connect jira-read without jira-admin.
| Split | Example | Why the window cares |
|---|---|---|
| Read vs write | jira-read / jira-write | Read jobs never see write schemas |
| Domain | billing vs scm vs chat | Host connects one domain per job type |
| Tenant | Per-customer server or per-token scope | Auth-scoped tools/list (current spec) can hide verbs; still filter the prompt |
| Frequency | Always-on get_issue vs deferred admin verbs | Keep the hot path in the prefix; defer the rest |
Description rules that actually cut tokens:
- One or two sentences: when to choose this tool, not how the SaaS works.
- Argument descriptions: units, id provenance, “never invent this id.”
- No pasted API docs. No HTML. No changelog.
- No duplicate paragraph that already lives in the system prompt.
- Enums for verbs the handler supports — not open strings that force a longer description to apologize.
- Each server has an owner who feels a page
- Read servers cannot call write handlers even if the model asks
- Descriptions reviewed for length the same way we review system prompts
-
tools/listorder is stable (current spec should; helps caches) - You did not “fix tokens” by renaming 90 tools into one omnibus
callwith a free-formactionstring — that moves the zoo into an argument the model will invent
A short schema on a write tool is still a write tool. Policy sits in the host. Length is not authorization.
Prompt cache cares about stability. The current MCP tools spec calls out deterministic list order because reshuffling the same tools still looks like a new prefix to a cache. Hosts should also keep allowlist membership stable across turns of one run. Adding Slack on turn 3 because the model “might need it” busts the cache and re-bills the new zoo.
| Change mid-run | Cache / window effect | Do this instead |
|---|---|---|
| Tool order shuffled | Miss; same tokens re-billed | Sort by name; pin the array |
| New server attached mid-job | Prefix grows; cache miss | New job type, new run |
listChanged → full re-dump | Miss plus possible overflow | Re-intersect the allowlist only |
| Deferred tool hydrated at end | Vendor-dependent; OpenAI documents end-of-window inject for tool search | Verify live docs; do not shuffle the static prefix |
Stable catalog, then small. Unstable catalog with a router is still a new prefix every turn if you are sloppy.
What guardrails belong in the host this week?
The host is the only place that can stop the dump. Servers will happily list everything they own. Models will happily read whatever you attach. You are the filter.
| Guardrail | Where it lives | Failure if skipped |
|---|---|---|
| Job-typed allowlist | Worker config | Zoo on every job |
| Policy gate on writes | Before tools/call | Smaller catalog, same double-send |
| Result cap / digest | Server + host | Tax two overflow |
Trace of the tools array | Observability | You cannot measure |
| Server connect allowlist | Deploy / desktop policy | Laptop config leaks to prod |
| Spec pin | Manifest | You built against a blog post, not the current MCP spec |
Anti-patterns (same week, do not ship these):
| Anti-pattern | What it does to the window |
|---|---|
| Enable every MCP server “for the demo” | Prefix death on turn 1 |
Drain all tools/list pages into the provider request | Pagination theater |
| Treat vendor tool-search as automatic in every host | Desktop may still dump; verify |
| Quote Anthropic’s 55K / 150K examples as your SLA | Wrong catalog, wrong number |
| Put the allowlist only in the system prompt | The model can ignore prose and still see the schemas |
| Return the SaaS key in a tool result “for flexibility” | Now the key lives in context and logs |
Unbounded list_* defaults | Catalog fix, result blow |
Week-scale checklist (skip the platform rewrite):
- Disconnect servers the golden job does not need
- Ship one job type with a hard allowlist and a trace
- Cap the one list tool that job calls
- Add wrong-tool to the eval for that job
- Write the number of prefix tokens on the design doc
- Leave router work for the week you can prove the static list lost
Eval the catalog the way you eval the job. A passing ticket lookup with the zoo still attached is not a catalog win.
| Eval case | Catalog fixture | Pass if |
|---|---|---|
| Status lookup | {get_issue} only | Calls get_issue; never create/transition |
| Status lookup + twins | {get_issue, create_issue, search_issues} | Still get_issue; wrong-tool = fail |
| Zoo control | Full connected set | Documents prefix tokens; may fail overflow — that is a finding |
| Result cap | {list_issues} with a fat fixture | Returns ≤ cap; model asks for next page if needed |
| Write job | {get_issue, add_comment} | Comment once; idempotency key from runtime |
- Catalog fixture is checked in next to the prompt fixture
- Zoo control run is labeled as a tax measurement, not a product demo
- CI fails if a read job’s fixture gains a write verb
- Prefix token count is stored as a metric, not a screenshot
Confirm MCP details you depend on — pagination, list caching, auth-scoped lists, outputSchema, session headers — against the current public spec. The 2026-07-28 line is what this post checked. Specs move. Your pin should too.
FAQ
Why did loading all my MCP tools blow the context window?
Because every attached tool schema is prompt tokens. The host copied the listed catalog into the model request, so the window filled with names, descriptions, and JSON Schema before your question ran. You do not have to call a tool for it to cost context. Count the tools array you actually sent.
How do I measure whether a smaller MCP catalog is working?
Serialize the provider tools payload and count tokens with the billed tokenizer, then track wrong-tool rate and overflows on one golden job. Compare that prefix to the same job after the allowlist. If prefix tokens drop and the job still passes, the catalog cut worked. Inspector tools/list size is the wrong metric.
What usually fails first when teams try this?
Turn-one context errors, or a completed turn that called a lookalike write tool because twins were all in the prompt. Next is turn-two overflow from an uncapped list result. Desktop configs that worked with three servers fail when the worker attaches twenty. Measure the request, not the settings UI.
How long does this take to show results?
A disconnect-and-allowlist pass on one job type is same-day work if you can export a trace. Description cleanup and read/write server splits take a few days. Routers and sandbox progressive disclosure take longer and should wait on a measured prefix. You should see token and wrong-tool movement on the first golden job, not after a platform rewrite.
What should I skip if I only have a week?
Skip building a custom embedding router, wrapping every helper in a new MCP server, and quoting vendor example token counts as your forecast. Disconnect unused servers, ship one job-typed allowlist, cap the list tool, and put prefix tokens on the trace. Leave defer_loading for the week the static list is still too fat.
When is this not worth doing yet?
When you do not have an agent in production, when the path is already a workflow, or when you cannot export the model request. Catalog surgery without a trace is fashion. If you cannot write pass/fail for the job, fix that first — when not to build an agent still applies. A two-tool native loop does not need a zoo policy.
CTA
Stop sending the zoo. Send the job.
What questions does this article answer?
- Why did loading all my MCP tools blow the context window?
- Because every attached tool schema is prompt tokens. The host copied the listed catalog into the model request, so the window filled with names, descriptions, and JSON Schema before your question ran. You do not have to call a tool for it to cost context. Count the `tools` array you actually sent.
- How do I measure whether a smaller MCP catalog is working?
- Serialize the provider `tools` payload and count tokens with the billed tokenizer, then track wrong-tool rate and overflows on one golden job. Compare that prefix to the same job after the allowlist. If prefix tokens drop and the job still passes, the catalog cut worked. Inspector `tools/list` size is the wrong metric.
- What usually fails first when teams try this?
- Turn-one context errors, or a completed turn that called a lookalike write tool because twins were all in the prompt. Next is turn-two overflow from an uncapped list result. Desktop configs that worked with three servers fail when the worker attaches twenty. Measure the request, not the settings UI.
- How long does this take to show results?
- A disconnect-and-allowlist pass on one job type is same-day work if you can export a trace. Description cleanup and read/write server splits take a few days. Routers and sandbox progressive disclosure take longer and should wait on a measured prefix. You should see token and wrong-tool movement on the first golden job, not after a platform rewrite.
- What should I skip if I only have a week?
- Skip building a custom embedding router, wrapping every helper in a new MCP server, and quoting vendor example token counts as your forecast. Disconnect unused servers, ship one job-typed allowlist, cap the list tool, and put prefix tokens on the trace. Leave `defer_loading` for the week the static list is still too fat.
- When is this not worth doing yet?
- When you do not have an agent in production, when the path is already a workflow, or when you cannot export the model request. Catalog surgery without a trace is fashion. If you cannot write pass/fail for the job, fix that first — [when not to build an agent](/blog/when-not-to-build-an-agent) still applies. A two-tool native loop does not need a zoo policy.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.