Spurlock Studios
Contact
Share LinkedIn X
Nested brass frames. Thesis: TOOL SCHEMAS AGENTS FOLLOW DESCRIPTIONS.

Agents invent arguments when your tool schema is vague. Treat JSON Schema as agent-facing UX: every property gets a description, finite sets become enums, required fields are honest, and the omnibus do_anything tool dies before it becomes privilege escalation. Schema quality is not documentation polish — it is how you stop silent production bugs.

This spoke belongs to the Agentic Systems Operating Manual. It owns the shape the model is allowed to emit. Write tools still need idempotent agent tool writes so a valid argument that times out does not double-send.

The short answer

  • Descriptions are the UX. Empty property descriptions force the model to guess units, formats, and when to omit fields.
  • Enums beat free text for status codes, channels, and action verbs the handler actually supports.
  • required must match reality. Optional-in-handler but required-in-schema (or the reverse) creates retry storms.
  • Strict mode grammar-constrains valid JSON — and rejects schemas outside the provider’s supported subset.
  • Split omnibus tools. One mega-tool with a free-form action string is a confused deputy with an API.
  • Measure argument accuracy on a golden set the same way you measure task pass rate.

What a schema the model will actually obey looks like

Vendors do not consume “full JSON Schema.” They consume a subset, then constrain generation against that subset. If your schema sits outside the subset, the request 400s before the model runs. If it sits inside but stays vague, the model emits legal JSON that your handler still cannot use.

The portable core across OpenAI parameters, Anthropic input_schema, Gemini function declarations, and MCP inputSchema:

KeywordWhat the model does with itWhat you still validate
typePicks a JSON typeCoerce vs reject ("30" vs 30)
descriptionDecides meaning and when to fillUnits, ID provenance
enumCollapses inventable stringsStale enums after an upstream change
requiredDecides which keys must appearHandler defaults vs schema defaults
additionalProperties: falseBlocks mystery keysNested objects you forgot to close

JSON Schema itself allows extra object keys unless you say otherwise. Models will use that slack. Close it.

Minimal production object (strict-safe: every property listed, optionals nullable):

{
  "type": "object",
  "additionalProperties": false,
  "properties": {
    "contact_id": {
      "type": "string",
      "description": "CRM contact UUID from search_contacts. Never invent an id."
    },
    "lifecycle_stage": {
      "type": "string",
      "enum": ["subscriber", "lead", "opportunity", "customer", "churned"],
      "description": "Exact lifecycle stage value accepted by the CRM. Use only these enums."
    },
    "note": {
      "type": ["string", "null"],
      "description": "Optional internal note, max ~500 chars. Null if no note."
    }
  },
  "required": ["contact_id", "lifecycle_stage", "note"]
}

That object is the contract. The system prompt is commentary.

Why empty property descriptions cause production bugs

The model sees property names. Names lie.

Schema smellWhat the model inventsProduction bug
amount with no descriptionDollars vs centsOff-by-100 charges
date as string, no formattomorrow, 02/03/26, unixHandler rejects the string → retry loop
status free stringclosed, Complete, doneDownstream enum reject
user vs user_idEmail stuffed into id field404 / wrong tenant
Nested options: {} emptyHallucinated keysStrict mode 400 or silent ignore

Anthropic’s define-tools page is blunt: detailed descriptions are “by far the most important factor in tool performance,” and they want 3–4 sentences on the tool itself plus a description on every parameter. Gemini’s function calling guide says the same thing in fewer words — be specific, give examples, put closed sets in enum.

Worked failure: a schedule_meeting tool exposed duration with no description. The model sent 30 (minutes) on Monday and "30m" on Tuesday after a prompt tweak. Half the calendar API calls 400’d; the agent retried until budget died. One sentence — "Duration in minutes as an integer, e.g. 30" — would have prevented the week of noise.

Description checklist for every write-path property:

  • Unit named (minutes, cents, ISO-8601 date, not “a number”)
  • Provenance named (from search_contacts, not “the id”)
  • Negative named (Never invent, Do not pass a display name)
  • Example that matches type (integer example for integer fields)

If you cannot write those four, the property is not ready for an agent.

What makes a tool description agent-facing UX

The top-level description answers three questions the model asks every turn:

  1. When should I call this (triggers)?
  2. When must I not call it (negative triggers)?
  3. What side effect happens if I do?

OpenAI’s function calling field table calls description “details on when and how to use the function.” That is trigger language, not a docstring restating the name.

Weak:

Updates a CRM record.

Stronger:

Update an existing CRM contact by contact_id. Call when the user confirmed a field change
(email, phone, lifecycle stage). Do not call to create contacts — use create_contact.
Do not call when the change is still a draft awaiting human approval.

Prescriptive when beats poetic what. Recent models that reach for tools conservatively especially need trigger language.

Description jobPut it hereDo not put it here
When / when-notTop-level descriptionBuried in a 2k-token system prompt only
Units and formatsPer-property descriptionA comment in the handler
Closed value setenumA prose list the model can “almost” match
Rare format the model keeps wronginput_examples (Anthropic) plus descriptionA sixth paragraph of the same sentence

Anthropic’s input_examples field is optional and schema-validated. Invalid examples 400 the request. Use them for nested objects and date formats. Do not use them as a substitute for enum.

Enums beat free text for finite sets

JSON Schema enum restricts a value to a fixed array. That is the whole point. If the handler only accepts five lifecycle stages, a free string is a bug you scheduled.

Field classFree string the model inventsEnum you should have shipped
CRM stageevangelist, Closed Won, hotValues the CRM API documents today
Channelsms, text, iMessageemail, sms if those are the two handlers
Verb on a writeupsert, patch, syncThe verbs the switch statement implements
TimezoneEST, eastern, New YorkIANA names you actually resolve

OpenAI’s function-calling guide says it directly: use enums and object structure to make invalid states unrepresentable. A toggle_light(on: bool, off: bool) tool is a schema that permits on=true, off=true. That is not a model failure. That is your schema.

Gemini’s function-declaration schema is a subset of OpenAPI, not full JSON Schema. Their docs tell you to set enum on a typed string. Teams that generate from Zod and omit "type": "string" next to enum have watched Gemini reject the declaration. If you generate, assert the provider-safe shape in CI — do not trust the default exporter.

Enum rules that hold in production:

  1. The array matches the handler switch / the upstream API, not the product glossary.
  2. You version the enum when the CRM adds a stage.
  3. Tool errors name the rejected value and list the allowed set so the model can correct once.
  4. You do not “helpfully” add synonyms (complete and done) unless both hit the same code path.

An enum that is two releases behind the CRM is how a correct model enters a retry loop.

Required fields that match the handler

required is a contract with two readers: the model and your function. If they disagree, you pay in retries.

Schema saysHandler doesWhat you see online
note required, type stringTreats missing note as fineModel invents filler notes
note omitted from requiredThrows if note is absentIntermittent 500s
strict: true, note not in requiredNever reachedProvider 400 on the request
note required + type: ["string","null"]Null means “no note”Portable optional

OpenAI strict mode requires every key in properties to appear in required, and every object to set additionalProperties: false. Optional-in-product becomes required-but-nullable. Chat Completions stays non-strict unless you set strict: true. Responses will try to normalize into strict and fall back to best-effort if the schema cannot be made compatible — look at the echoed strict: false when that happens.

Anthropic is looser on optional keys in the non-strict path and still wants required to name what the tool cannot run without. Do not copy an OpenAI-strict schema onto Anthropic (or the reverse) without a transform.

Procedure when a field feels optional:

  1. Ask whether the handler can run with the key absent.
  2. If yes and you are on OpenAI strict: keep the key, allow null, list it in required.
  3. If yes and you are on a non-strict provider: omit it from required, describe the default in the property text.
  4. If no: required, non-null, description that says so.

Dishonest required arrays are the quietest source of “the agent keeps calling the same tool.”

additionalProperties: false is not optional in strict mode

JSON Schema’s default is permissive. Extra keys are valid unless you close the object. Understanding JSON Schema says this in the first paragraph on additionalProperties.

OpenAI Structured Outputs goes further: additionalProperties: false must be set on every object or strict: true is rejected. Nested objects count. Arrays of objects count.

ObjectClosed?What leaks if you forget
Root parameters / input_schemaMust beMystery top-level keys
Nested addressMust beaddress.line3, address.notes
items object inside an arrayMust bePer-row junk the handler ignores
MCP no-arg toolRecommended closed empty objectA model that still sends { "ok": true }

MCP’s 2026-07-28 tools spec is explicit for tools with no parameters: prefer { "type": "object", "additionalProperties": false } so the only legal arguments object is {}. { "type": "object" } alone still accepts arbitrary keys.

additionalProperties: false is not authorization. A closed schema can still update the wrong contact if contact_id was invented. Close the object, then validate IDs against a search you already ran.

JSON Schema keywords agents ignore or providers reject

This is the portability trap. Full JSON Schema has pattern, minimum, maximum, format, $ref, allOf, if/then. Provider tool surfaces do not honor the same set, and they do not fail the same way.

KeywordOpenAI strictAnthropic toolsGemini declarationsWhat to do instead
enumYesYesYes, on stringsPrefer this
descriptionYesYesYesRequired in your style guide
patternOften accepted, not a generation guaranteeMay reject in strictSubset / ignoreValidate in the handler; describe the pattern in text
minimum / maximumSameSamePartialHandler bounds + description (1–100)
$ref / deep $defsLimited / rejected in strictLimitedLimitedFlatten or inline
format: date-timeHint at bestHintDocumented for some string formatsDescription + handler

OpenAI’s structured-outputs page is the source: if you turn on strict: true with an unsupported schema, you get an error. That is a gift. The worse outcome is a keyword the API accepts and the sampler does not enforce — your CI is green, production still sees "30m".

Portable rule: describe bounds, enforce bounds in the handler, return a field-level error. Do not bet the write path on pattern.

Strict mode — when it helps and when it hurts

Helps: write tools where invalid JSON arguments must never reach the handler; high-volume agents where “almost valid” burns retries.

Hurts / fails closed:

  • Schema uses unsupported keywords → API 400 before the model runs
  • You needed a loosely typed escape hatch for rare admin ops
  • You generated schemas from rich Pydantic/Zod models without a “strict-safe” transform

OpenAI documents the two hard constraints: every object closed, every property required. They recommend enabling strict. They also list limitations — some JSON Schema features are unsupported, and certain multi-function paths on fine-tuned models disable it. Read the current function calling page before you assume last quarter’s exporter still compiles.

Procedure when enabling strict:

  1. Freeze the schema subset your provider documents
  2. Turn on strict: true (OpenAI tools / Anthropic tool definitions)
  3. Run the golden set; collect schema compiler errors separately from model errors
  4. Keep handler validation anyway — strict is not authorization

Strict mode is grammar. A grammatically perfect merge_contacts call can still be a policy violation. Gate the side effect in the harness, not in the JSON Schema compiler.

Killing the omnibus tool

An omnibus tool looks like this:

{
  "name": "crm",
  "description": "Do anything in the CRM",
  "input_schema": {
    "type": "object",
    "properties": {
      "action": { "type": "string" },
      "payload": { "type": "object" }
    },
    "required": ["action", "payload"]
  }
}

Why it fails in production:

  1. No enum on action → invents verbs the handler does not implement
  2. payload is a bag → no property descriptions, no required fields
  3. Privilege escalation → one allowlisted tool name unlocks every CRM write
  4. Eval blindness → you cannot score “correct tool choice” when there is only one tool

That is the confused deputy problem with a friendlier name. The model is not the deputy you intended to empower. The tool name is. If crm is allowlisted, merge_contacts rode in for free.

Split by side-effect class and audience:

Split toolSide effectWho may call
search_contactsReadAgent
update_contact_stageWrite (reversible)Agent + policy
merge_contactsWrite (hard)Human approval only
export_contacts_csvRead / bulkDeny in prod agent

Fewer, sharper tools beat one god-tool every time. If you need composition, put it in your code — not in a free-form action string.

Split test: if two calls share a name but not a blast radius, they are two tools.

MCP tool schemas vs provider-native function schemas

Same JSON Schema idea; different wrappers, field names, and dialects.

SurfaceParameters keyWrapperDialect notes
OpenAI function toolsparameters{ "type": "function", "name", "description", "parameters", "strict?" }Strict subset; closed objects
Anthropic toolsinput_schemaTop-level { name, description, input_schema, strict?, input_examples? }JSON Schema with documented limits
Gemini function declarationsparametersOpenAPI-subset object on the declarationenum wants an explicit string type
MCP toolsinputSchemaServer-advertised tool; host translates for the modelDefault JSON Schema 2020-12; root type must be object

MCP’s tools spec also adds outputSchema and structuredContent. If you declare an output schema, the server must return structured results that match it. That is a second contract. Most teams still ship inputSchema only and then wonder why hosts cannot type the result.

Consequences:

  • You cannot paste an MCP tool record into OpenAI’s tools array without translation
  • MCP hosts discover tools at runtime — giant catalogs blow context (token tax)
  • Execution location differs: native function calling usually runs in your harness; MCP often runs in a server process with its own credentials
  • MCP says tools are model-controlled and still tells implementers there should be a human who can deny a call

Practical pattern: keep one canonical schema (Zod/Pydantic/JSON), generate provider adapters, and generate MCP inputSchema from the same source. Do not hand-maintain three drifting copies.

MCP tools/list is paginated and cacheable. The spec asks servers to return a deterministic order so prompt-cache hits survive. A 80-tool dump into every turn is not “being complete.” It is paying rent on tools the model will mis-pick.

Should schemas be generated from Pydantic / Zod?

Yes — with a strict-safe export path.

Pydantic model_json_schema() emits a 2020-12-ish dict, including enum from Python Enums and $defs for nested models. That is a good start and a bad finish. The default export is built for validators, not for OpenAI strict or Gemini’s OpenAPI subset. Zod’s JSON Schema export has the same job to do: strip keywords the provider rejects, close every object, force descriptions, and flatten $ref when strict cannot follow them.

Checklist:

  • Generator emits additionalProperties: false on every object
  • Every field has a description string (enforce in CI)
  • Enums for closed sets, not open strings
  • Target flag matches the provider (openAiStrict / Anthropic-safe / Gemini-safe)
  • Nullable optionals match the provider’s optional story
  • Golden-set fixtures assert on generated schema hashes so silent regen diffs fail CI

Generated schemas without descriptions are how teams ship empty UX at scale. A missing Field(description=...) is a production bug, not a style nit.

If the exporter cannot satisfy a provider, fail the build. Do not “just turn strict off” on the write path to make CI green.

How many tools before the catalog becomes noise

Schema quality per tool dies when the catalog is a junk drawer. Gemini’s function-calling docs tell you to keep the active set around 10–20 tools. That is a practical ceiling, not a law, and it matches what we see on agent loops: past a few dozen similarly named writes, tool-choice accuracy falls before argument accuracy does.

Catalog smellWhat the model doesFix
update_record and update_contactPicks the vaguer oneRename to the noun the user said
12 tools whose descriptions start “Use this to…”Skims and guessesLead with the trigger, not the filler
MCP server that advertises admin + prodCalls the admin verbSeparate servers / filter tools/list by audience
Read and write twins with the same stemWrites when it meant to readAsymmetric names: search_ vs update_

Active-set procedure:

  1. Inventory tools by side-effect class (read / reversible write / irreversible write / admin)
  2. Expose only the class this run is allowed to use
  3. Keep the rest registered in the host, not in the model context
  4. Re-score tool-choice on the golden set after every add

A perfect schema on the 40th lookalike tool still loses to a smaller, sharper list.

Measuring tool-call argument accuracy

Do not wait for “task failed” to learn the schema is wrong.

Golden-set columns that matter:

ColumnExample
expected_toolupdate_contact_stage
expected_args{ "contact_id": "…", "lifecycle_stage": "customer" }
forbidden_toolsmerge_contacts
notesAmbiguous user text on purpose

Score separately:

  1. Tool choice accuracy — right tool name
  2. Argument exact-match / schema-valid — right shape
  3. Semantic arg match — same meaning after normalization (dates, phones)

A schema change that lifts argument exact-match while task pass rate stays flat still shipped value — fewer retries, lower cost, fewer weird writes. Do not hide that lift inside a single “task success” number.

Eval hygiene:

  • Compiler errors (schema rejected by the provider) are a separate bucket from model errors
  • Each write tool has at least one “almost right” case (30m vs 30, done vs complete)
  • Forbidden tools are scored, not only the happy path
  • You re-run after exporter or CRM enum changes, not only after prompt edits

If you cannot name the last time argument exact-match moved, you are flying on anecdotes.

First schema review checklist

  • Tool name is a verb_noun the handler implements (update_contact_stage, not crm)
  • Top-level description has triggers and negative triggers
  • Every property has a non-empty description
  • Finite sets are enum
  • required matches handler reality (or required-plus-nullable under strict)
  • additionalProperties: false on objects
  • No omnibus payload: object without inner schema
  • Side-effect class documented for the policy gate
  • Example args in docs or input_examples if your provider supports them
  • Golden cases cover the three failure modes you fear most
  • Write tools do not let the model invent idempotency keys — those are minted in the runtime, per idempotent agent tool writes

If a write tool fails that list, it does not ship. Read tools can ship with two holes. Money and email tools cannot.

Failure mode: schema drift after the CRM upgrade

What broke: CRM added lifecycle_stage values; the tool schema enum stayed frozen; the model correctly wanted evangelist; strict mode / handler rejected; agent looped.

What it cost: retries until the budget died, a contact stuck in opportunity, and a human who “fixed” it by widening the schema to a free string — which reopened invented stages.

Fix path:

  1. Schema version in the tool registry
  2. Contract test against a live or stubbed CRM enum endpoint
  3. Golden case for each new stage before deploy
  4. Alert on repeated invalid_enum tool errors online
  5. Refuse the “just make it a string” hotfix on write tools

Schemas are APIs. Version them.

Second failure we keep seeing: the exporter adds a new Pydantic field, CI does not hash the schema, strict mode starts 400ing on Monday, and the on-call blames the model. Hash the generated schema. Fail the build when it changes without a golden-set update.

What the handler must still reject

A schema the model obeys is necessary and not sufficient. The handler is the second reader.

CheckWhy the schema cannot do it alone
ID exists in this tenantThe model can emit a well-typed UUID it invented
Caller is allowed to write this recordSchema has no notion of authz
Amount within a business capmaximum may not be enforced by the sampler
Idempotency key is the runtime’s, not the model’sA model-minted key is a new key on every retry
Enum still matches upstreamSchema copy is a snapshot

Return errors the model can fix: field name, rejected value, allowed set or format. An opaque 500 Internal teaches the agent to retry the same bad call. A lifecycle_stage must be one of: subscriber, lead, opportunity, customer, churned teaches it to pick customer.

Never execute on a 400 you could have explained. The cheapest token is the one you do not spend on a third identical write.

Pilot minimum

In a Spurlock $1,500 · 5-day agentic pilot: inventory tools and kill omnibus entries; rewrite descriptions + enums on the write path; add a 20–40 case argument-accuracy slice; return handler validation errors the model can fix (not opaque 500s). Zero empty property descriptions on tools that send email or move money.

FAQ

When should I split one “do_anything” tool into many?

As soon as one tool name can perform more than one side-effect class, or when action/payload is free-form. Split by verb and risk: reads vs reversible writes vs irreversible writes. Omnibus tools hide privilege and make evals meaningless.

Do input_examples help more than longer descriptions?

They help different jobs. Descriptions win for when/when-not triggers and units. Examples win for formats the model keeps wrong. For write tools, use both — but put closed sets in enum before you write a novel.

Strict mode — when does it hurt?

When your schema uses keywords the provider’s strict compiler rejects, or when you still need a rare loosely typed admin escape hatch. It also does not replace authorization. Use strict on high-volume write tools after the schema is inside the supported subset; keep handler validation.

How do MCP tool schemas differ from provider function schemas?

MCP advertises inputSchema; OpenAI expects parameters inside a function tool object; Anthropic expects input_schema; Gemini uses an OpenAPI-subset declaration. The JSON Schema core can be shared, but the wrappers differ — translate from one canonical source. MCP also shifts discovery and execution to servers, which changes credential boundaries and context cost.

Should schemas be generated from Pydantic/Zod?

Yes, if CI enforces descriptions, enums, additionalProperties: false, and a provider-safe export. Generated schemas without those constraints just industrialize empty UX. Hash schemas in golden-set CI so silent regen cannot drift production.

What’s a good first schema review checklist?

Name, triggers, negative triggers, every property described, enums for finite sets, honest required, no omnibus payload bags, side-effect class for the gate, and golden cases for the scary paths. If a write tool fails that list, it does not ship.

CTA

Want schemas that stop inventing arguments on a real job in five days? Start at /agentic or /contact?intent=agentic-pilot.

FAQ

What questions does this article answer?

When should I split one “do_anything” tool into many?
As soon as one tool name can perform more than one side-effect class, or when `action`/`payload` is free-form. Split by verb and risk: reads vs reversible writes vs irreversible writes. Omnibus tools hide privilege and make evals meaningless.
Do input_examples help more than longer descriptions?
They help different jobs. Descriptions win for when/when-not triggers and units. Examples win for formats the model keeps wrong. For write tools, use both — but put closed sets in `enum` before you write a novel.
Strict mode — when does it hurt?
When your schema uses keywords the provider’s strict compiler rejects, or when you still need a rare loosely typed admin escape hatch. It also does not replace authorization. Use strict on high-volume write tools after the schema is inside the supported subset; keep handler validation.
How do MCP tool schemas differ from provider function schemas?
MCP advertises `inputSchema`; OpenAI expects `parameters` inside a function tool object; Anthropic expects `input_schema`; Gemini uses an OpenAPI-subset declaration. The JSON Schema core can be shared, but the wrappers differ — translate from one canonical source. MCP also shifts discovery and execution to servers, which changes credential boundaries and context cost.
Should schemas be generated from Pydantic/Zod?
Yes, if CI enforces descriptions, enums, `additionalProperties: false`, and a provider-safe export. Generated schemas without those constraints just industrialize empty UX. Hash schemas in golden-set CI so silent regen cannot drift production.
What’s a good first schema review checklist?
Name, triggers, negative triggers, every property described, enums for finite sets, honest `required`, no omnibus payload bags, side-effect class for the gate, and golden cases for the scary paths. If a write tool fails that list, it does not ship.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot