Spurlock Studios
Contact
Share LinkedIn X
Nested brass frames. Thesis: DID LOADING ALL MCP TOOLS.

Loading every MCP tool blew the context window because the host stuffed each tool’s name, description, and JSON Schema into the model request. That tax is billed before the user sentence. You did not run forty tools. You advertised forty tools. The model had to read the zoo in order to pick one cage.

This is not the planning-loop tax of later turns. That bill is turns × a growing transcript. This bill is the prefix: schemas sitting in context on turn one. It belongs under the Agentic Systems Operating Manual. If the job is a known lookup with one write, you may not want an agent at all — that brake lives in when not to build an agent.

The short answer

  • Schemas are context. OpenAI’s function-calling guide is blunt: callable definitions count against the context limit and are billed as input tokens.
  • MCP discovery ≠ the prompt. As of the 2026-07-28 tools spec, tools/list tells the client what exists. The host still chooses what the model sees this turn.
  • Load a subset or a router. Job-typed allowlist first. Vendor defer_loading / tool search, or a search_tools meta-tool, when the measured prefix still crowds the task.
  • Do not quote a fake constant. Token cost depends on description length, nested JSON Schema, and whether the host forwards outputSchema and other metadata. Count your serialized payload.
  • Results are tax two. A fat ticket dump on turn one will blow the window even after you shrink the catalog.

Why did a one-line task still blow the window?

Because the user message was never the bulk of the request. The bulk was the tool catalog the host attached so the model could “have access.” Access is not free. The model has to read every schema you sent in order to decide which one to call.

Anthropic’s code execution with MCP (Nov 4, 2025) names the two overload patterns: tool definitions loaded upfront, and intermediate results passed through the model. Their write-up describes agents connected to hundreds or thousands of tools across dozens of servers. Treat that as vendor framing of the failure, not a census of your stack.

What you thought happenedWhat the request actually containedWho pays
“I asked one question”System prompt + every attached tool schema + the questionInput tokens, turn 1
“I have not called GitHub yet”GitHub’s full inputSchema set is already in the promptSame
“MCP is just a connector”The host copied tools/list into the provider tools arrayHost policy, not the spec
“The window is 200k, we are fine”Prefix of tens of thousands of tokens leaves little room for the task, history, and resultsYou, when the completion truncates or the model picks the wrong twin tool

A one-line task with an 80-tool prefix is a small question taped to a large dictionary. The dictionary is the bill.

What actually entered the model request?

Pin the layers. People say “MCP used all my context” when they mean “my host forwarded every listed tool into the provider request.” Those are different objects.

LayerWhat it holdsDoes it occupy the model window?
MCP server processHandlers, secrets, real side effectsNo
tools/list RPCName, description, inputSchema, optional outputSchema, icons, annotationsOnly if the host copies it onward
Host allowlist / routerThe subset this job may seeIndirectly: it decides the next row
Provider tools / input_schema arrayThe schemas the model is allowed to call nowYes. This is the prefix tax.
Tool result on the next turnJSON / text the handler returnedYes. This is tax two.

Procedure for a blown window:

  1. Export the raw model request (not the MCP inspector view).
  2. Isolate the tools / functions array from messages.
  3. Serialize that array exactly as the SDK sent it.
  4. Count tokens with the same tokenizer the provider bills, on that payload alone.
  5. Compare that number to the user message. If the catalog dwarfs the question, you found the leak.

If you cannot export the request, you are guessing. Guessing is how “always 500 tokens per tool” folklore gets invented. Kill the folklore. Count the bytes you actually sent.

Native function calling pays the same prefix tax. MCP did not invent “schemas in the prompt.” MCP made it easy to attach twenty servers you would never have typed by hand. The protocol is the delivery truck. The host is who loaded the truck.

How the window “blows”What operators seeWhat it is
Hard overflowProvider 400 / context-length error on turn 1Prefix + messages exceeded the model limit
Silent squeezeRequest succeeds; task, history, or results get truncated by the hostYou “fixed” the error by dropping the work
Quality collapseCompletes, wrong tool, ignores the questionCatalog crowding; twins in the prompt
Cache thrashStable system prompt, but every turn misses the cacheTool array order or membership changed

“Blew the window” is any of those four. Only the first looks like a red error. The other three still waste money and miss the ticket.

Is tools/list the same as the catalog the model sees?

No. Mixing those two is the usual architecture error.

The 2026-07-28 tools spec (confirm this line on the current public docs before you promise a client a feature) says servers that declare tools must answer tools/list. The set may be empty, may change over time, must not vary per-connection as a side effect of other requests, and may vary by the authorization on the request. Listing supports pagination and caching. Servers should return a deterministic order so clients can cache and so prompt-cache hit rates improve when tools are included in model context.

That last clause is the tell. The spec knows tools often land in the model prompt. It does not require the host to include every listed tool.

JobConsumerHealthy behavior
tools/list (paged)Host / gatewayLearn what the server could run
Auth-scoped listSameHide tools the token cannot use
Per-run allowlistHostChoose verbs for this job
Provider tool arrayThe modelOnly the allowlist, or a router plus a handful of always-on tools

Pagination does not save the window by itself. A client that drains every nextCursor and then forwards the union to the model has a tidy RPC and a ruined prompt. Cursor opacity, server-chosen page size, missing limit= in the protocol — those are wire facts. They are not a catalog policy.

If your “fix” was “the server paginates tools/list now,” you fixed the JSON-RPC payload. You have not fixed the model request until the host stops stuffing the union.

listChanged (current tools spec) means the server may notify that the list moved. Useful for the host cache. Dangerous if the host reacts by re-dumping a new union into the next completion and breaking the prompt cache. Treat a list-changed notification as “refresh the allowlist intersection,” not “send everything again.”

Hosts that also dump resources and prompts into the model are running a second zoo. MCP paginates resources/list and prompts/list for the same reason it paginates tools. The window still dies if the host inlines resource bodies or every prompt template on turn one.

MCP surfaceList method (current spec)Model-window rule
Toolstools/listAttach the job subset, not the union
Resourcesresources/listFetch by id when the job needs the blob
Promptsprompts/listPick the template for this job; do not paste the library
Templatesresources/templates/listSame as resources: discover ≠ inline

Confirm those method names on the live spec if you are pinning an older line. Do not assume 2025 SSE-era client code still matches.

Why are schemas expensive when nobody called a tool?

Because the model cannot call what it cannot see, and seeing is reading. A tool the handler never ran still occupied context on every turn you attached it.

Cost drivers that show up in real catalogs (ranges, not constants):

DriverWhy tokens climbWhat to do
Long descriptionProse is tokens. Essays in the description field are prompt bloat.Write “when to pick this tool,” not the handbook.
Nested JSON SchemaEvery properties key, enum, and $ref expansion is text.Keep arguments the handler actually reads.
Duplicate twinsjira_create_issue vs jira_create_ticket vs create_issueOne verb per side effect.
outputSchema forwardedOptional on current MCP tools spec; hosts that copy it add more prompt.Keep output schema for validation in the host if you need it; do not assume the model must eat it.
Examples inside the schemaFew-shot in JSON is still tokens every turn.Put examples in evals, or in a deferred doc the router can fetch.
Icons / titles / extra metadataHarmless as URLs; painful if a host inlines blobs.Do not inline image bytes into the tool array.

Same tool, two description styles. Count these strings with your tokenizer. Do not reuse whatever number you get as a constant for GitHub’s server.

Fat (handbook in the schema — avoid this):

{
  "name": "get_issue",
  "description": "This tool retrieves a Jira issue from the configured cloud site. Jira is Atlassian’s issue tracker. Use this whenever the user mentions a ticket, incident, bug, story, or epic. The issue key looks like PROJ-123. Never guess. The response includes the full rendered HTML of the description, all comments, all attachments metadata, watchers, sprint, and custom fields. If you are unsure whether to use this or search_issues or create_issue or transition_issue, consider the user’s intent carefully and prefer this tool for lookups.",
  "inputSchema": {
    "type": "object",
    "properties": {
      "issue_key": { "type": "string", "description": "The issue key." },
      "expand": { "type": "string", "description": "Optional expand query. You can pass many comma-separated values." },
      "include_html": { "type": "boolean" },
      "include_comments": { "type": "boolean" },
      "include_attachments": { "type": "boolean" }
    }
  }
}

Thin (when to pick it — prefer this):

{
  "name": "get_issue",
  "description": "Fetch one Jira issue by key (e.g. INC-1842). Use for status lookups. Do not use to create or transition issues.",
  "inputSchema": {
    "type": "object",
    "additionalProperties": false,
    "required": ["issue_key"],
    "properties": {
      "issue_key": {
        "type": "string",
        "description": "Existing Jira key from the user or a prior search. Never invent a key."
      }
    }
  }
}

The fat version is still one tool. It still belongs in the prefix on every turn you attach it. Multiply that style across forty verbs and you will not need a vendor anecdote to know why the window died.

OpenAI’s function-calling page says that if you hit token limits, limit the number of functions loaded up front, shorten descriptions, or use tool search so deferred tools load only when needed. That is the vendor admitting the zoo is expensive. It is not a license to invent “tool X always costs N.”

Anthropic’s advanced tool use post (Nov 24, 2025) published a worked example of one five-server mix: GitHub 35 tools (~26K tokens), Slack 11 (~21K), Sentry 5 (~3K), Grafana 5 (~3K), Splunk 2 (~2K) — about 58 tools and ~55K tokens before the conversation starts, in their catalog. Jira in that post is ~17K alone. They also report seeing tool definitions consume 134K tokens before optimization. Those are measurements of specific servers on a specific date. Recite them as Anthropic’s example. Do not paste 55K into a customer deck as if it were your GitHub server.

How do desktop and IDE hosts dump the zoo?

Most people meet this bug in a desktop or IDE host, not in a custom worker. You tick twelve MCP servers in settings because they might be useful. The host connects all of them, lists every tool, and — unless it has a defer path — attaches the union to the next completion.

Custom workers repeat the same mistake with more confidence: await Promise.all(servers.map(listTools)), then chat.completions.create({ tools: everything }).

Host shapeTypical dumpWhat “subset” means here
Desktop / IDE MCP panelEvery enabled server, every listed tool, every sessionDisable servers you are not using this hour; prefer hosts that defer
Overnight agent workerEvery server in the deploy manifest, every job typeFilter by job type in code before the model call
Multi-tenant gatewayEvery tool the caller’s token can seeScope the token and the prompt; auth-scoped tools/list is not enough if you still forward the list
Eval harnessGolden-set tools plus leftovers from a local configPin the catalog in the eval fixture

Checklist for the host you actually run:

  • I can name which servers are connected in production, not just in my laptop config
  • I can name which of those servers this job type is allowed to see
  • I can show the tools array on a trace for that job type
  • I have a way to disable a noisy server without a full redeploy
  • I have not copied my desktop MCP config into the worker “so prod matches my IDE”

Desktop convenience is not a production catalog. If the worker’s first request includes Slack, Sentry, Grafana, and a filesystem server so it can “be helpful,” you paid the zoo tax on a job that needed one ticket tool.

Failure mode: forty servers, one ticket lookup

What breaks: A support agent is asked “What is the status of INC-1842?” The host has GitHub, Jira, Slack, Confluence, Sentry, a SQL server, a browser server, and a filesystem server connected. The model spends the prefix reading create-issue, close-issue, list-channels, run-query, and write-file. It then either (a) overflows, (b) picks a lookalike write tool, or (c) answers from a stale description because the right get_issue schema was a needle in the stack.

What it costs: Input tokens on every turn, worse tool selection, and a smaller remainder for the actual ticket body. If the first successful call returns the full issue JSON including comments, attachments, and HTML, tax two finishes the job the schemas started.

What you do instead:

  1. Job type support.status_lookup → allowlist {jira.get_issue} (read).
  2. If the run must comment, that is a second job type with {jira.get_issue, jira.add_comment} and a policy gate.
  3. Writes still use runtime idempotency keys — see idempotent agent tool writes. A smaller catalog does not make a double-post safe.
  4. Cap the get_issue result: identifiers, status, assignee, last comment. Not the attachment bytes.
SymptomLikely layerFirst fix
Request 400 / context-length error on turn 1Prefix schemasDrop servers; allowlist
Completes but calls create_issue for a status questionTwin tools in the promptKill the write verb for this job type
Turn 1 fine, turn 2 overflowsTool resultTruncate / paginate / digest
Cache miss every turnCatalog reshuffledDeterministic tool order (spec should); stable allowlist
“It worked in the IDE”Desktop had three servers; worker has twentyPin catalogs per environment

This is the failure we keep seeing after 20,000+ hours on agentic systems: the demo used two tools. Production inherited the founder’s MCP panel.

How do I measure the tax without inventing a constant?

You measure the serialized catalog the host sends, plus the result payloads the loop appends. You do not average a blog-post token number across vendors.

MetricHow to take itPass heuristic (yours to set)
Prefix tool tokensTokenizer on the tools array onlySmall vs the model’s context; leave room for the task
User+system tokensSame request, messages onlyShould dominate the question, not lose to the zoo
Result tokens / turnEach tool payload as sent back to the modelCap; Anthropic notes Claude Code defaults a 25k-token tool-response cap — that is their product default, not a spec limit
Turns until overflowTrace of a golden jobJob completes with remainder
Wrong-tool rateEval on lookalike namesDrops after you remove twins

Numbered measurement procedure:

  1. Pick one golden job (the ticket question, not “be a general assistant”).
  2. Run it with the current zoo. Export the request.
  3. Count prefix tool tokens. Record wrong-tool and overflow.
  4. Cut to the allowlist you believe the job needs. Repeat the same golden job.
  5. If prefix is still huge, shorten descriptions and split servers. Count again.
  6. Only then consider a router / defer_loading. Routers add a search turn; you should know the static allowlist was not enough.

Checklist so the number stays honest:

  • Same model pin for before/after
  • Same tokenizer as billing, not a guess from a third-party widget
  • Catalog taken from the request, not from tools/list in the inspector
  • Result caps measured separately from schema tokens
  • You wrote the allowlist down so next week’s demo cannot silently reattach Slack

If someone quotes “MCP tools always cost N tokens,” ask for the serialized payload. No payload, no number.

Worksheet to paste into the design doc (fill with your traces):

FieldZoo runAllowlist runNotes
Date / model pinSame pin both columns
Servers connectedNames, not “all”
Tools in provider arrayCount of schemas
Prefix tool tokensTokenizer = billed
User+system tokens
Result tokens turn 1After the first call
Wrong-tool? (Y/N + name)
Overflow? (Y/N + turn)
Job pass/failSame golden prompt

Empty cells mean you did not measure. Empty cells are not “probably fine.”

What does a job-sized subset look like?

A job-sized subset is the verbs this job is allowed to use, not the company’s platform catalog. Three to eight tools is a common production shape for a single job type. That is a design default, not a law. A lookup with one read tool is better than eight. A run that truly needs twelve related reads should still not include an unrelated write.

Job typeServers connectedTools the model seesTools the server still owns
support.status_lookupJiraget_issuecreate, transition, attach, search
support.commentJiraget_issue, add_commentcreate, delete
release.notesGitHublist_merged_prsevery other GitHub verb
ops.pagePagerDutyget_incident, ack_incidentcreate, delete, list-everything

Allowlist procedure:

  1. Name the job type in the runner (string, not a vibe).
  2. Map job type → server set (usually one).
  3. Map job type → tool name set (usually a handful).
  4. Intersect with what tools/list returned so a renamed server fails closed.
  5. Pass that intersection to the model. Log the names. Fail the run if the intersection is empty.
  • Every production job type has a written allowlist
  • Write verbs are absent from read jobs
  • Omnibus do_anything / call_api tools are not on the list
  • The allowlist is in the worker, not in a prompt paragraph the model can ignore
  • Changing the allowlist is a reviewed config change

The server can keep a large implementation surface. The model should not.

When do I want a router instead of a static allowlist?

When one worker honestly serves many job types and a static list per type is either huge or constantly wrong. A router is a search step: the model (or the host) retrieves a handful of schemas, then calls them. It is extra latency. Pay it when you have measured that the static prefix is the problem.

PatternWhat the model sees firstWho searchesWhen it fits
Static allowlist3–8 full schemasYou, at job-type routing timeDefault. Most pilots.
Host pre-filterSchemas for this tenant × jobYour gateway, no extra model callMulti-tenant, known verbs
OpenAI tool_search + defer_loadingServer/namespace name+description; schemas laterProvider or your client-executed searchLarge declared inventory; confirm current model support on OpenAI’s tool-search guide
Anthropic Tool Search Tool + defer_loadingSearch tool + any non-deferred toolsClaude searches, then full schemas loadMCP-heavy hosts; confirm the current Claude tool-use docs, not only the Nov 2025 announcement
Code / filesystem progressive disclosureDirectory of servers, read schemas on demandThe model reads files or a search_tools toolYou already run a sandbox; Anthropic’s MCP code-execution post reports a worked example dropping 150,000 → 2,000 tokens in that setup — their numbers, not yours

OpenAI’s tool-search guide (as of the public page cited here) recommends grouping deferred functions into namespaces or MCP servers with clear high-level descriptions, and treating “fewer than 10 functions per namespace” as a best-practice for their search quality. Confirm that guidance on the live page. Do not treat “10” as an MCP spec limit. There is none.

Anthropic’s same Nov 2025 post suggests considering Tool Search when tool definitions exceed ~10K tokens, when selection accuracy is already bad, when multiple MCP servers are in play, or when 10+ tools are available — and says it is less useful when the library is small or every tool is used every session. Again: their product heuristic. Your measurement from the previous section outranks it.

Router checklist:

  • Static allowlist already lost on a measured prefix or wrong-tool rate
  • Search corpus is names and descriptions that a retrieval step can rank
  • Write tools cannot be retrieved into a read job without a policy gate
  • You budgeted one extra turn for search
  • Prompt cache: deferred tools should not scramble a stable prefix (OpenAI documents injecting discovered tools at the end of the window for that reason — verify current behavior)

If you cannot name the job types, a router will retrieve a slightly different zoo every turn. That is not a fix. That is a slot machine with schemas.

Hosted vs client-executed search (vendor names differ; the split is the same):

  1. Hosted: You declare the inventory (functions, namespaces, or MCP servers) with defer_loading where the vendor supports it, add the search tool, and let the API choose what to hydrate. Confirm the current request shape on OpenAI’s tool-search page or Anthropic’s current tool-use docs before you copy a year-old gist.
  2. Client-executed: The model emits a search call. Your gateway queries an index you own (job-type allowlist, BM25, embeddings — you pick). You return the matching full schemas. Then the model calls a real tool.
  3. Policy after retrieval: A retrieved delete_repo still must fail the gate on a read job. Search is not authorization.
  4. Log both hops: search query, names returned, names actually called. If search returns twelve write twins, your index descriptions are lying.

A search_tools meta-tool you write yourself is valid. Keep it small: query in, a handful of {name, description, inputSchema} out, hard cap on hits. If it returns the whole catalog “just in case,” you built a router that dumps the zoo on purpose.

Results are a second window tax

Shrinking the catalog and then returning a 10,000-row sheet into the next message is how teams “fix MCP” and still overflow. Definitions are tax one. Results are tax two. Anthropic’s MCP code-execution post uses a 2-hour sales transcript as an example of ~50,000 additional tokens when the full document flows through the model twice (read, then write). That is an illustration of a fat payload, not a constant for every transcript.

Result shapeWindow effectPrefer
Full document / sheet / logCan exceed the remainder after schemasFilter in the tool; return a digest + id
Unbounded list_*Page 1 is 20 rows; page “all” is the archiveDefault limit; cursor; refuse limit > cap
Nested objects with HTMLTags are tokensPlain fields the next tool needs
Error tracebacksStack frames crowd the retryOne-line, actionable error the model can correct
Secrets in resultsTranscript + logs now hold the keyNever return secrets; the model will see them

Procedure for every list-shaped MCP tool:

  1. Required limit with a small default (you pick; Anthropic’s agent-tools post argues for pagination, range, filter, and truncation with sensible defaults).
  2. Opaque cursor; never invent one on the client.
  3. Explicit truncation flag so the model knows to ask for the next page.
  4. Hard cap in the server, not only in the prompt.
  5. For binaries and long docs, return a handle the next tool can fetch — do not paste the file into the assistant message.

The MCP pagination spec covers tools/list, resources/list, prompts/list, and templates. It does not paginate tools/call results for you. Result envelopes are your server’s job. If a list tool returns the whole table, that is a handler bug, not an MCP omission.

Writes that survive this filter still need idempotent retries. Lookup-plus-comment is two side-effect classes. Keys stay in the runtime.

When is a workflow enough instead of an 80-tool agent?

When you can name the path. “Get INC-1842, format status, post to Slack” is a workflow with two HTTP nodes. Putting those two calls behind MCP so an agent can also create issues, delete files, and run SQL is how a status bot inherits a company-wide blast radius and a blown window.

Read when not to build an agent before you keep the loop. The catalog tax is one more reason that page exists: agents pay for optionality. Workflows do not attach 80 unused schemas.

SignalWorkflowAgent with a tiny catalogAgent with a router
Path known, IDs in the triggerYesNoNo
Path varies, pass/fail writable, verbs ≤ ~8Maybe a hybridYesNot yet
Many job types, one worker, measured prefix painSplit workersSplit job types firstThen router
“Might need every SaaS someday”NoNoNo

Decision list:

  1. If the trigger already contains the issue id, do not search the universe.
  2. If a human would run the same three clicks every time, encode the clicks.
  3. If the model must choose among a few reads, give it those reads — not the write twins.
  4. If you cannot write pass/fail, you are not ready for a catalog discussion. You are not ready for an agent.

Optionality you do not use is still billed. That is the whole post.

How should I split servers and shrink descriptions?

Split by blast radius and job, not by vendor logo. One “company MCP server” with 90 tools is a gift to context bloat and to confused-deputy writes. Several small servers let a host connect jira-read without jira-admin.

SplitExampleWhy the window cares
Read vs writejira-read / jira-writeRead jobs never see write schemas
Domainbilling vs scm vs chatHost connects one domain per job type
TenantPer-customer server or per-token scopeAuth-scoped tools/list (current spec) can hide verbs; still filter the prompt
FrequencyAlways-on get_issue vs deferred admin verbsKeep the hot path in the prefix; defer the rest

Description rules that actually cut tokens:

  1. One or two sentences: when to choose this tool, not how the SaaS works.
  2. Argument descriptions: units, id provenance, “never invent this id.”
  3. No pasted API docs. No HTML. No changelog.
  4. No duplicate paragraph that already lives in the system prompt.
  5. Enums for verbs the handler supports — not open strings that force a longer description to apologize.
  • Each server has an owner who feels a page
  • Read servers cannot call write handlers even if the model asks
  • Descriptions reviewed for length the same way we review system prompts
  • tools/list order is stable (current spec should; helps caches)
  • You did not “fix tokens” by renaming 90 tools into one omnibus call with a free-form action string — that moves the zoo into an argument the model will invent

A short schema on a write tool is still a write tool. Policy sits in the host. Length is not authorization.

Prompt cache cares about stability. The current MCP tools spec calls out deterministic list order because reshuffling the same tools still looks like a new prefix to a cache. Hosts should also keep allowlist membership stable across turns of one run. Adding Slack on turn 3 because the model “might need it” busts the cache and re-bills the new zoo.

Change mid-runCache / window effectDo this instead
Tool order shuffledMiss; same tokens re-billedSort by name; pin the array
New server attached mid-jobPrefix grows; cache missNew job type, new run
listChanged → full re-dumpMiss plus possible overflowRe-intersect the allowlist only
Deferred tool hydrated at endVendor-dependent; OpenAI documents end-of-window inject for tool searchVerify live docs; do not shuffle the static prefix

Stable catalog, then small. Unstable catalog with a router is still a new prefix every turn if you are sloppy.

What guardrails belong in the host this week?

The host is the only place that can stop the dump. Servers will happily list everything they own. Models will happily read whatever you attach. You are the filter.

GuardrailWhere it livesFailure if skipped
Job-typed allowlistWorker configZoo on every job
Policy gate on writesBefore tools/callSmaller catalog, same double-send
Result cap / digestServer + hostTax two overflow
Trace of the tools arrayObservabilityYou cannot measure
Server connect allowlistDeploy / desktop policyLaptop config leaks to prod
Spec pinManifestYou built against a blog post, not the current MCP spec

Anti-patterns (same week, do not ship these):

Anti-patternWhat it does to the window
Enable every MCP server “for the demo”Prefix death on turn 1
Drain all tools/list pages into the provider requestPagination theater
Treat vendor tool-search as automatic in every hostDesktop may still dump; verify
Quote Anthropic’s 55K / 150K examples as your SLAWrong catalog, wrong number
Put the allowlist only in the system promptThe model can ignore prose and still see the schemas
Return the SaaS key in a tool result “for flexibility”Now the key lives in context and logs
Unbounded list_* defaultsCatalog fix, result blow

Week-scale checklist (skip the platform rewrite):

  • Disconnect servers the golden job does not need
  • Ship one job type with a hard allowlist and a trace
  • Cap the one list tool that job calls
  • Add wrong-tool to the eval for that job
  • Write the number of prefix tokens on the design doc
  • Leave router work for the week you can prove the static list lost

Eval the catalog the way you eval the job. A passing ticket lookup with the zoo still attached is not a catalog win.

Eval caseCatalog fixturePass if
Status lookup{get_issue} onlyCalls get_issue; never create/transition
Status lookup + twins{get_issue, create_issue, search_issues}Still get_issue; wrong-tool = fail
Zoo controlFull connected setDocuments prefix tokens; may fail overflow — that is a finding
Result cap{list_issues} with a fat fixtureReturns ≤ cap; model asks for next page if needed
Write job{get_issue, add_comment}Comment once; idempotency key from runtime
  • Catalog fixture is checked in next to the prompt fixture
  • Zoo control run is labeled as a tax measurement, not a product demo
  • CI fails if a read job’s fixture gains a write verb
  • Prefix token count is stored as a metric, not a screenshot

Confirm MCP details you depend on — pagination, list caching, auth-scoped lists, outputSchema, session headers — against the current public spec. The 2026-07-28 line is what this post checked. Specs move. Your pin should too.

FAQ

Why did loading all my MCP tools blow the context window?

Because every attached tool schema is prompt tokens. The host copied the listed catalog into the model request, so the window filled with names, descriptions, and JSON Schema before your question ran. You do not have to call a tool for it to cost context. Count the tools array you actually sent.

How do I measure whether a smaller MCP catalog is working?

Serialize the provider tools payload and count tokens with the billed tokenizer, then track wrong-tool rate and overflows on one golden job. Compare that prefix to the same job after the allowlist. If prefix tokens drop and the job still passes, the catalog cut worked. Inspector tools/list size is the wrong metric.

What usually fails first when teams try this?

Turn-one context errors, or a completed turn that called a lookalike write tool because twins were all in the prompt. Next is turn-two overflow from an uncapped list result. Desktop configs that worked with three servers fail when the worker attaches twenty. Measure the request, not the settings UI.

How long does this take to show results?

A disconnect-and-allowlist pass on one job type is same-day work if you can export a trace. Description cleanup and read/write server splits take a few days. Routers and sandbox progressive disclosure take longer and should wait on a measured prefix. You should see token and wrong-tool movement on the first golden job, not after a platform rewrite.

What should I skip if I only have a week?

Skip building a custom embedding router, wrapping every helper in a new MCP server, and quoting vendor example token counts as your forecast. Disconnect unused servers, ship one job-typed allowlist, cap the list tool, and put prefix tokens on the trace. Leave defer_loading for the week the static list is still too fat.

When is this not worth doing yet?

When you do not have an agent in production, when the path is already a workflow, or when you cannot export the model request. Catalog surgery without a trace is fashion. If you cannot write pass/fail for the job, fix that first — when not to build an agent still applies. A two-tool native loop does not need a zoo policy.

CTA

Stop sending the zoo. Send the job.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

Why did loading all my MCP tools blow the context window?
Because every attached tool schema is prompt tokens. The host copied the listed catalog into the model request, so the window filled with names, descriptions, and JSON Schema before your question ran. You do not have to call a tool for it to cost context. Count the `tools` array you actually sent.
How do I measure whether a smaller MCP catalog is working?
Serialize the provider `tools` payload and count tokens with the billed tokenizer, then track wrong-tool rate and overflows on one golden job. Compare that prefix to the same job after the allowlist. If prefix tokens drop and the job still passes, the catalog cut worked. Inspector `tools/list` size is the wrong metric.
What usually fails first when teams try this?
Turn-one context errors, or a completed turn that called a lookalike write tool because twins were all in the prompt. Next is turn-two overflow from an uncapped list result. Desktop configs that worked with three servers fail when the worker attaches twenty. Measure the request, not the settings UI.
How long does this take to show results?
A disconnect-and-allowlist pass on one job type is same-day work if you can export a trace. Description cleanup and read/write server splits take a few days. Routers and sandbox progressive disclosure take longer and should wait on a measured prefix. You should see token and wrong-tool movement on the first golden job, not after a platform rewrite.
What should I skip if I only have a week?
Skip building a custom embedding router, wrapping every helper in a new MCP server, and quoting vendor example token counts as your forecast. Disconnect unused servers, ship one job-typed allowlist, cap the list tool, and put prefix tokens on the trace. Leave `defer_loading` for the week the static list is still too fat.
When is this not worth doing yet?
When you do not have an agent in production, when the path is already a workflow, or when you cannot export the model request. Catalog surgery without a trace is fashion. If you cannot write pass/fail for the job, fix that first — [when not to build an agent](/blog/when-not-to-build-an-agent) still applies. A two-tool native loop does not need a zoo policy.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot