Tool Schemas Agents Follow: Descriptions, Enums, and Killing the Omnibus Tool
Agents invent arguments when schemas are vague. Write JSON Schema like agent UX—enums, required fields, property descriptions—and kill the do_anything tool.
William Spurlock Founder — Spurlock Studios Updated 16 MIN
Agents invent arguments when your tool schema is vague. Treat JSON Schema as agent-facing UX: every property gets a description, finite sets become enums, required fields are honest, and the omnibus do_anything tool dies before it becomes privilege escalation. Schema quality is not documentation polish — it is how you stop silent production bugs.
This spoke belongs to the Agentic Systems Operating Manual. It owns the shape the model is allowed to emit. Write tools still need idempotent agent tool writes so a valid argument that times out does not double-send.
The short answer
- Descriptions are the UX. Empty property descriptions force the model to guess units, formats, and when to omit fields.
- Enums beat free text for status codes, channels, and action verbs the handler actually supports.
requiredmust match reality. Optional-in-handler but required-in-schema (or the reverse) creates retry storms.- Strict mode grammar-constrains valid JSON — and rejects schemas outside the provider’s supported subset.
- Split omnibus tools. One mega-tool with a free-form
actionstring is a confused deputy with an API. - Measure argument accuracy on a golden set the same way you measure task pass rate.
What a schema the model will actually obey looks like
Vendors do not consume “full JSON Schema.” They consume a subset, then constrain generation against that subset. If your schema sits outside the subset, the request 400s before the model runs. If it sits inside but stays vague, the model emits legal JSON that your handler still cannot use.
The portable core across OpenAI parameters, Anthropic input_schema, Gemini function declarations, and MCP inputSchema:
| Keyword | What the model does with it | What you still validate |
|---|---|---|
type | Picks a JSON type | Coerce vs reject ("30" vs 30) |
description | Decides meaning and when to fill | Units, ID provenance |
enum | Collapses inventable strings | Stale enums after an upstream change |
required | Decides which keys must appear | Handler defaults vs schema defaults |
additionalProperties: false | Blocks mystery keys | Nested objects you forgot to close |
JSON Schema itself allows extra object keys unless you say otherwise. Models will use that slack. Close it.
Minimal production object (strict-safe: every property listed, optionals nullable):
{
"type": "object",
"additionalProperties": false,
"properties": {
"contact_id": {
"type": "string",
"description": "CRM contact UUID from search_contacts. Never invent an id."
},
"lifecycle_stage": {
"type": "string",
"enum": ["subscriber", "lead", "opportunity", "customer", "churned"],
"description": "Exact lifecycle stage value accepted by the CRM. Use only these enums."
},
"note": {
"type": ["string", "null"],
"description": "Optional internal note, max ~500 chars. Null if no note."
}
},
"required": ["contact_id", "lifecycle_stage", "note"]
}
That object is the contract. The system prompt is commentary.
Why empty property descriptions cause production bugs
The model sees property names. Names lie.
| Schema smell | What the model invents | Production bug |
|---|---|---|
amount with no description | Dollars vs cents | Off-by-100 charges |
date as string, no format | tomorrow, 02/03/26, unix | Handler rejects the string → retry loop |
status free string | closed, Complete, done | Downstream enum reject |
user vs user_id | Email stuffed into id field | 404 / wrong tenant |
Nested options: {} empty | Hallucinated keys | Strict mode 400 or silent ignore |
Anthropic’s define-tools page is blunt: detailed descriptions are “by far the most important factor in tool performance,” and they want 3–4 sentences on the tool itself plus a description on every parameter. Gemini’s function calling guide says the same thing in fewer words — be specific, give examples, put closed sets in enum.
Worked failure: a schedule_meeting tool exposed duration with no description. The model sent 30 (minutes) on Monday and "30m" on Tuesday after a prompt tweak. Half the calendar API calls 400’d; the agent retried until budget died. One sentence — "Duration in minutes as an integer, e.g. 30" — would have prevented the week of noise.
Description checklist for every write-path property:
- Unit named (
minutes,cents,ISO-8601 date, not “a number”) - Provenance named (
from search_contacts, not “the id”) - Negative named (
Never invent,Do not pass a display name) - Example that matches
type(integer example for integer fields)
If you cannot write those four, the property is not ready for an agent.
What makes a tool description agent-facing UX
The top-level description answers three questions the model asks every turn:
- When should I call this (triggers)?
- When must I not call it (negative triggers)?
- What side effect happens if I do?
OpenAI’s function calling field table calls description “details on when and how to use the function.” That is trigger language, not a docstring restating the name.
Weak:
Updates a CRM record.
Stronger:
Update an existing CRM contact by contact_id. Call when the user confirmed a field change
(email, phone, lifecycle stage). Do not call to create contacts — use create_contact.
Do not call when the change is still a draft awaiting human approval.
Prescriptive when beats poetic what. Recent models that reach for tools conservatively especially need trigger language.
| Description job | Put it here | Do not put it here |
|---|---|---|
| When / when-not | Top-level description | Buried in a 2k-token system prompt only |
| Units and formats | Per-property description | A comment in the handler |
| Closed value set | enum | A prose list the model can “almost” match |
| Rare format the model keeps wrong | input_examples (Anthropic) plus description | A sixth paragraph of the same sentence |
Anthropic’s input_examples field is optional and schema-validated. Invalid examples 400 the request. Use them for nested objects and date formats. Do not use them as a substitute for enum.
Enums beat free text for finite sets
JSON Schema enum restricts a value to a fixed array. That is the whole point. If the handler only accepts five lifecycle stages, a free string is a bug you scheduled.
| Field class | Free string the model invents | Enum you should have shipped |
|---|---|---|
| CRM stage | evangelist, Closed Won, hot | Values the CRM API documents today |
| Channel | sms, text, iMessage | email, sms if those are the two handlers |
| Verb on a write | upsert, patch, sync | The verbs the switch statement implements |
| Timezone | EST, eastern, New York | IANA names you actually resolve |
OpenAI’s function-calling guide says it directly: use enums and object structure to make invalid states unrepresentable. A toggle_light(on: bool, off: bool) tool is a schema that permits on=true, off=true. That is not a model failure. That is your schema.
Gemini’s function-declaration schema is a subset of OpenAPI, not full JSON Schema. Their docs tell you to set enum on a typed string. Teams that generate from Zod and omit "type": "string" next to enum have watched Gemini reject the declaration. If you generate, assert the provider-safe shape in CI — do not trust the default exporter.
Enum rules that hold in production:
- The array matches the handler
switch/ the upstream API, not the product glossary. - You version the enum when the CRM adds a stage.
- Tool errors name the rejected value and list the allowed set so the model can correct once.
- You do not “helpfully” add synonyms (
completeanddone) unless both hit the same code path.
An enum that is two releases behind the CRM is how a correct model enters a retry loop.
Required fields that match the handler
required is a contract with two readers: the model and your function. If they disagree, you pay in retries.
| Schema says | Handler does | What you see online |
|---|---|---|
note required, type string | Treats missing note as fine | Model invents filler notes |
note omitted from required | Throws if note is absent | Intermittent 500s |
strict: true, note not in required | Never reached | Provider 400 on the request |
note required + type: ["string","null"] | Null means “no note” | Portable optional |
OpenAI strict mode requires every key in properties to appear in required, and every object to set additionalProperties: false. Optional-in-product becomes required-but-nullable. Chat Completions stays non-strict unless you set strict: true. Responses will try to normalize into strict and fall back to best-effort if the schema cannot be made compatible — look at the echoed strict: false when that happens.
Anthropic is looser on optional keys in the non-strict path and still wants required to name what the tool cannot run without. Do not copy an OpenAI-strict schema onto Anthropic (or the reverse) without a transform.
Procedure when a field feels optional:
- Ask whether the handler can run with the key absent.
- If yes and you are on OpenAI strict: keep the key, allow
null, list it inrequired. - If yes and you are on a non-strict provider: omit it from
required, describe the default in the property text. - If no: required, non-null, description that says so.
Dishonest required arrays are the quietest source of “the agent keeps calling the same tool.”
additionalProperties: false is not optional in strict mode
JSON Schema’s default is permissive. Extra keys are valid unless you close the object. Understanding JSON Schema says this in the first paragraph on additionalProperties.
OpenAI Structured Outputs goes further: additionalProperties: false must be set on every object or strict: true is rejected. Nested objects count. Arrays of objects count.
| Object | Closed? | What leaks if you forget |
|---|---|---|
Root parameters / input_schema | Must be | Mystery top-level keys |
Nested address | Must be | address.line3, address.notes |
items object inside an array | Must be | Per-row junk the handler ignores |
| MCP no-arg tool | Recommended closed empty object | A model that still sends { "ok": true } |
MCP’s 2026-07-28 tools spec is explicit for tools with no parameters: prefer { "type": "object", "additionalProperties": false } so the only legal arguments object is {}. { "type": "object" } alone still accepts arbitrary keys.
additionalProperties: false is not authorization. A closed schema can still update the wrong contact if contact_id was invented. Close the object, then validate IDs against a search you already ran.
JSON Schema keywords agents ignore or providers reject
This is the portability trap. Full JSON Schema has pattern, minimum, maximum, format, $ref, allOf, if/then. Provider tool surfaces do not honor the same set, and they do not fail the same way.
| Keyword | OpenAI strict | Anthropic tools | Gemini declarations | What to do instead |
|---|---|---|---|---|
enum | Yes | Yes | Yes, on strings | Prefer this |
description | Yes | Yes | Yes | Required in your style guide |
pattern | Often accepted, not a generation guarantee | May reject in strict | Subset / ignore | Validate in the handler; describe the pattern in text |
minimum / maximum | Same | Same | Partial | Handler bounds + description (1–100) |
$ref / deep $defs | Limited / rejected in strict | Limited | Limited | Flatten or inline |
format: date-time | Hint at best | Hint | Documented for some string formats | Description + handler |
OpenAI’s structured-outputs page is the source: if you turn on strict: true with an unsupported schema, you get an error. That is a gift. The worse outcome is a keyword the API accepts and the sampler does not enforce — your CI is green, production still sees "30m".
Portable rule: describe bounds, enforce bounds in the handler, return a field-level error. Do not bet the write path on pattern.
Strict mode — when it helps and when it hurts
Helps: write tools where invalid JSON arguments must never reach the handler; high-volume agents where “almost valid” burns retries.
Hurts / fails closed:
- Schema uses unsupported keywords → API 400 before the model runs
- You needed a loosely typed escape hatch for rare admin ops
- You generated schemas from rich Pydantic/Zod models without a “strict-safe” transform
OpenAI documents the two hard constraints: every object closed, every property required. They recommend enabling strict. They also list limitations — some JSON Schema features are unsupported, and certain multi-function paths on fine-tuned models disable it. Read the current function calling page before you assume last quarter’s exporter still compiles.
Procedure when enabling strict:
- Freeze the schema subset your provider documents
- Turn on
strict: true(OpenAI tools / Anthropic tool definitions) - Run the golden set; collect schema compiler errors separately from model errors
- Keep handler validation anyway — strict is not authorization
Strict mode is grammar. A grammatically perfect merge_contacts call can still be a policy violation. Gate the side effect in the harness, not in the JSON Schema compiler.
Killing the omnibus tool
An omnibus tool looks like this:
{
"name": "crm",
"description": "Do anything in the CRM",
"input_schema": {
"type": "object",
"properties": {
"action": { "type": "string" },
"payload": { "type": "object" }
},
"required": ["action", "payload"]
}
}
Why it fails in production:
- No enum on
action→ invents verbs the handler does not implement payloadis a bag → no property descriptions, no required fields- Privilege escalation → one allowlisted tool name unlocks every CRM write
- Eval blindness → you cannot score “correct tool choice” when there is only one tool
That is the confused deputy problem with a friendlier name. The model is not the deputy you intended to empower. The tool name is. If crm is allowlisted, merge_contacts rode in for free.
Split by side-effect class and audience:
| Split tool | Side effect | Who may call |
|---|---|---|
search_contacts | Read | Agent |
update_contact_stage | Write (reversible) | Agent + policy |
merge_contacts | Write (hard) | Human approval only |
export_contacts_csv | Read / bulk | Deny in prod agent |
Fewer, sharper tools beat one god-tool every time. If you need composition, put it in your code — not in a free-form action string.
Split test: if two calls share a name but not a blast radius, they are two tools.
MCP tool schemas vs provider-native function schemas
Same JSON Schema idea; different wrappers, field names, and dialects.
| Surface | Parameters key | Wrapper | Dialect notes |
|---|---|---|---|
| OpenAI function tools | parameters | { "type": "function", "name", "description", "parameters", "strict?" } | Strict subset; closed objects |
| Anthropic tools | input_schema | Top-level { name, description, input_schema, strict?, input_examples? } | JSON Schema with documented limits |
| Gemini function declarations | parameters | OpenAPI-subset object on the declaration | enum wants an explicit string type |
| MCP tools | inputSchema | Server-advertised tool; host translates for the model | Default JSON Schema 2020-12; root type must be object |
MCP’s tools spec also adds outputSchema and structuredContent. If you declare an output schema, the server must return structured results that match it. That is a second contract. Most teams still ship inputSchema only and then wonder why hosts cannot type the result.
Consequences:
- You cannot paste an MCP tool record into OpenAI’s
toolsarray without translation - MCP hosts discover tools at runtime — giant catalogs blow context (token tax)
- Execution location differs: native function calling usually runs in your harness; MCP often runs in a server process with its own credentials
- MCP says tools are model-controlled and still tells implementers there should be a human who can deny a call
Practical pattern: keep one canonical schema (Zod/Pydantic/JSON), generate provider adapters, and generate MCP inputSchema from the same source. Do not hand-maintain three drifting copies.
MCP tools/list is paginated and cacheable. The spec asks servers to return a deterministic order so prompt-cache hits survive. A 80-tool dump into every turn is not “being complete.” It is paying rent on tools the model will mis-pick.
Should schemas be generated from Pydantic / Zod?
Yes — with a strict-safe export path.
Pydantic model_json_schema() emits a 2020-12-ish dict, including enum from Python Enums and $defs for nested models. That is a good start and a bad finish. The default export is built for validators, not for OpenAI strict or Gemini’s OpenAPI subset. Zod’s JSON Schema export has the same job to do: strip keywords the provider rejects, close every object, force descriptions, and flatten $ref when strict cannot follow them.
Checklist:
- Generator emits
additionalProperties: falseon every object - Every field has a description string (enforce in CI)
- Enums for closed sets, not open strings
- Target flag matches the provider (
openAiStrict/ Anthropic-safe / Gemini-safe) - Nullable optionals match the provider’s optional story
- Golden-set fixtures assert on generated schema hashes so silent regen diffs fail CI
Generated schemas without descriptions are how teams ship empty UX at scale. A missing Field(description=...) is a production bug, not a style nit.
If the exporter cannot satisfy a provider, fail the build. Do not “just turn strict off” on the write path to make CI green.
How many tools before the catalog becomes noise
Schema quality per tool dies when the catalog is a junk drawer. Gemini’s function-calling docs tell you to keep the active set around 10–20 tools. That is a practical ceiling, not a law, and it matches what we see on agent loops: past a few dozen similarly named writes, tool-choice accuracy falls before argument accuracy does.
| Catalog smell | What the model does | Fix |
|---|---|---|
update_record and update_contact | Picks the vaguer one | Rename to the noun the user said |
| 12 tools whose descriptions start “Use this to…” | Skims and guesses | Lead with the trigger, not the filler |
| MCP server that advertises admin + prod | Calls the admin verb | Separate servers / filter tools/list by audience |
| Read and write twins with the same stem | Writes when it meant to read | Asymmetric names: search_ vs update_ |
Active-set procedure:
- Inventory tools by side-effect class (read / reversible write / irreversible write / admin)
- Expose only the class this run is allowed to use
- Keep the rest registered in the host, not in the model context
- Re-score tool-choice on the golden set after every add
A perfect schema on the 40th lookalike tool still loses to a smaller, sharper list.
Measuring tool-call argument accuracy
Do not wait for “task failed” to learn the schema is wrong.
Golden-set columns that matter:
| Column | Example |
|---|---|
expected_tool | update_contact_stage |
expected_args | { "contact_id": "…", "lifecycle_stage": "customer" } |
forbidden_tools | merge_contacts |
notes | Ambiguous user text on purpose |
Score separately:
- Tool choice accuracy — right tool name
- Argument exact-match / schema-valid — right shape
- Semantic arg match — same meaning after normalization (dates, phones)
A schema change that lifts argument exact-match while task pass rate stays flat still shipped value — fewer retries, lower cost, fewer weird writes. Do not hide that lift inside a single “task success” number.
Eval hygiene:
- Compiler errors (schema rejected by the provider) are a separate bucket from model errors
- Each write tool has at least one “almost right” case (
30mvs30,donevscomplete) - Forbidden tools are scored, not only the happy path
- You re-run after exporter or CRM enum changes, not only after prompt edits
If you cannot name the last time argument exact-match moved, you are flying on anecdotes.
First schema review checklist
- Tool name is a verb_noun the handler implements (
update_contact_stage, notcrm) - Top-level description has triggers and negative triggers
- Every property has a non-empty description
- Finite sets are
enum -
requiredmatches handler reality (or required-plus-nullable under strict) -
additionalProperties: falseon objects - No omnibus
payload: objectwithout inner schema - Side-effect class documented for the policy gate
- Example args in docs or
input_examplesif your provider supports them - Golden cases cover the three failure modes you fear most
- Write tools do not let the model invent idempotency keys — those are minted in the runtime, per idempotent agent tool writes
If a write tool fails that list, it does not ship. Read tools can ship with two holes. Money and email tools cannot.
Failure mode: schema drift after the CRM upgrade
What broke: CRM added lifecycle_stage values; the tool schema enum stayed frozen; the model correctly wanted evangelist; strict mode / handler rejected; agent looped.
What it cost: retries until the budget died, a contact stuck in opportunity, and a human who “fixed” it by widening the schema to a free string — which reopened invented stages.
Fix path:
- Schema version in the tool registry
- Contract test against a live or stubbed CRM enum endpoint
- Golden case for each new stage before deploy
- Alert on repeated
invalid_enumtool errors online - Refuse the “just make it a string” hotfix on write tools
Schemas are APIs. Version them.
Second failure we keep seeing: the exporter adds a new Pydantic field, CI does not hash the schema, strict mode starts 400ing on Monday, and the on-call blames the model. Hash the generated schema. Fail the build when it changes without a golden-set update.
What the handler must still reject
A schema the model obeys is necessary and not sufficient. The handler is the second reader.
| Check | Why the schema cannot do it alone |
|---|---|
| ID exists in this tenant | The model can emit a well-typed UUID it invented |
| Caller is allowed to write this record | Schema has no notion of authz |
| Amount within a business cap | maximum may not be enforced by the sampler |
| Idempotency key is the runtime’s, not the model’s | A model-minted key is a new key on every retry |
| Enum still matches upstream | Schema copy is a snapshot |
Return errors the model can fix: field name, rejected value, allowed set or format. An opaque 500 Internal teaches the agent to retry the same bad call. A lifecycle_stage must be one of: subscriber, lead, opportunity, customer, churned teaches it to pick customer.
Never execute on a 400 you could have explained. The cheapest token is the one you do not spend on a third identical write.
Pilot minimum
In a Spurlock $1,500 · 5-day agentic pilot: inventory tools and kill omnibus entries; rewrite descriptions + enums on the write path; add a 20–40 case argument-accuracy slice; return handler validation errors the model can fix (not opaque 500s). Zero empty property descriptions on tools that send email or move money.
FAQ
When should I split one “do_anything” tool into many?
As soon as one tool name can perform more than one side-effect class, or when action/payload is free-form. Split by verb and risk: reads vs reversible writes vs irreversible writes. Omnibus tools hide privilege and make evals meaningless.
Do input_examples help more than longer descriptions?
They help different jobs. Descriptions win for when/when-not triggers and units. Examples win for formats the model keeps wrong. For write tools, use both — but put closed sets in enum before you write a novel.
Strict mode — when does it hurt?
When your schema uses keywords the provider’s strict compiler rejects, or when you still need a rare loosely typed admin escape hatch. It also does not replace authorization. Use strict on high-volume write tools after the schema is inside the supported subset; keep handler validation.
How do MCP tool schemas differ from provider function schemas?
MCP advertises inputSchema; OpenAI expects parameters inside a function tool object; Anthropic expects input_schema; Gemini uses an OpenAPI-subset declaration. The JSON Schema core can be shared, but the wrappers differ — translate from one canonical source. MCP also shifts discovery and execution to servers, which changes credential boundaries and context cost.
Should schemas be generated from Pydantic/Zod?
Yes, if CI enforces descriptions, enums, additionalProperties: false, and a provider-safe export. Generated schemas without those constraints just industrialize empty UX. Hash schemas in golden-set CI so silent regen cannot drift production.
What’s a good first schema review checklist?
Name, triggers, negative triggers, every property described, enums for finite sets, honest required, no omnibus payload bags, side-effect class for the gate, and golden cases for the scary paths. If a write tool fails that list, it does not ship.
CTA
Want schemas that stop inventing arguments on a real job in five days? Start at /agentic or /contact?intent=agentic-pilot.
What questions does this article answer?
- When should I split one “do_anything” tool into many?
- As soon as one tool name can perform more than one side-effect class, or when `action`/`payload` is free-form. Split by verb and risk: reads vs reversible writes vs irreversible writes. Omnibus tools hide privilege and make evals meaningless.
- Do input_examples help more than longer descriptions?
- They help different jobs. Descriptions win for when/when-not triggers and units. Examples win for formats the model keeps wrong. For write tools, use both — but put closed sets in `enum` before you write a novel.
- Strict mode — when does it hurt?
- When your schema uses keywords the provider’s strict compiler rejects, or when you still need a rare loosely typed admin escape hatch. It also does not replace authorization. Use strict on high-volume write tools after the schema is inside the supported subset; keep handler validation.
- How do MCP tool schemas differ from provider function schemas?
- MCP advertises `inputSchema`; OpenAI expects `parameters` inside a function tool object; Anthropic expects `input_schema`; Gemini uses an OpenAPI-subset declaration. The JSON Schema core can be shared, but the wrappers differ — translate from one canonical source. MCP also shifts discovery and execution to servers, which changes credential boundaries and context cost.
- Should schemas be generated from Pydantic/Zod?
- Yes, if CI enforces descriptions, enums, `additionalProperties: false`, and a provider-safe export. Generated schemas without those constraints just industrialize empty UX. Hash schemas in golden-set CI so silent regen cannot drift production.
- What’s a good first schema review checklist?
- Name, triggers, negative triggers, every property described, enums for finite sets, honest `required`, no omnibus payload bags, side-effect class for the gate, and golden cases for the scary paths. If a write tool fails that list, it does not ship.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.