How do I write good tool schemas for AI agents
Write tool schemas as agent UX: honest required fields, enums for closed sets, descriptions that constrain, and one non-overlapping tool per side effect.
William Spurlock Founder — Spurlock Studios 22 MIN
You write good tool schemas for AI agents by treating JSON Schema as agent-facing UX: one tool per side effect, an honest required array, enums for every closed set, and property descriptions that name units, provenance, and forbidden substitutes. Vague names plus empty descriptions are how the model invents arguments. Overlapping tools are how it picks the wrong write. This is a writing procedure, not a longer system prompt.
This spoke sits under the Agentic Systems Operating Manual. It is the how-to. The diagnosis of vague schemas, strict-mode grammar, and the omnibus do_anything tool lives in tool schemas agents follow. Score the result the way you build the evaluator before the agent: a golden set that can fail tool choice and argument shape separately.
The short answer
- Inventory side effects first. JSON is last.
- One
verb_nounper blast radius. If two tools can satisfy the same user sentence, they overlap — rename, add a negative trigger, or delete one. requiredmatches the handler. Under OpenAI strict mode, optional-in-product is required-plus-null.- Finite values are
enum, copied from the switch, not from the glossary. - Every property description constrains: unit, where the id came from, what never to pass, one example that matches
type.
| You are writing | Done looks like | Not done |
|---|---|---|
| Name | hold_order, not orders | Two stems that both mean “fix the order” |
| Top-level description | When / when-not / side effect / sibling tool | A restatement of the function name |
required | Keys the handler cannot run without | “Nice to have” fields the model will invent |
| Closed set | enum equal to the API | A comma list inside the description |
| Property text | Unit + provenance + never-invent | Empty string, or the name repeated |
How do I write good tool schemas for AI agents?
You write them in a fixed order so the model never sees a tool whose handler and schema disagree. Do not start in the JSON editor.
Writing order:
- List every side effect this job is allowed to perform (
read,write_reversible,write_irreversible,exfil_risk). - Run the overlap test on the names you want to ship (next section). Zero overlaps, or you stop.
- Name each tool
verb_nounthe handler actually implements. - Write the top-level
descriptionas trigger language — OpenAI’s function calling field table callsdescription“details on when and how to use the function.” - Open the handler. Mark every argument it cannot default. Those keys are
required. - For each finite set, copy the handler
switch/ upstream enum intoenum. Gemini’s function calling guide says to put closed sets inenuminstead of only describing them. - Write a constraining description on every property. Anthropic’s define tools page calls detailed descriptions “by far the most important factor in tool performance” and wants several sentences on the tool plus a description on each parameter.
- Close every object with
additionalProperties: false. JSON Schema treats extra keys as valid unless you say otherwise. - Write the handler error the model should see on a bad value: field name, rejected value, allowed set.
- Add golden cases: one happy path, one almost-right arg, one overlap trap.
| Step | Artifact | Reader |
|---|---|---|
| 1–2 | Side-effect table + overlap sheet | You |
| 3–4 | name + top-level description | Model (tool choice) |
| 5–8 | required, enum, property text, closed objects | Model (argument shape) |
| 9–10 | Error strings + golden set | Handler + evaluator |
If you skip to step 7 because “the types are obvious,” you will ship a legal schema the model still mis-calls. Types without triggers are a dictionary, not a tool.
What do I write down before any JSON Schema?
You write a side-effect card per tool. The card is the source of truth. The schema is a projection of the card into JSON Schema.
| Card field | Example | Why it exists |
|---|---|---|
name | hold_order | Tool-choice string |
| Side-effect class | write_reversible | Policy gate, not the model |
| Blast radius | One order, reversible by release_hold | Overlap test |
| Trigger | Operator confirmed a hold; order_id already from search_orders | Top-level description |
| Negative trigger | Do not use to cancel, refund, or change address | Sibling tools |
| Handler cannot-run-without | order_id, hold_reason | required |
| Closed sets | hold_reason ∈ four API values | enum |
| Id provenance | search_orders hit this turn | Property description |
| Forbidden substitutes | Confirmation number, email, “the last order” | Property description |
Card procedure:
- Sit with the operator who already does this job by hand. Ask what they click, what they refuse, and which id they copy from which screen.
- Fill the card in their words. If they cannot name a negative trigger, the tool is still an omnibus in disguise.
- Only then open the schema file.
- Diff the card against the handler signature. The card loses. The handler wins.
Interview questions that actually fill the card:
| Ask | You are extracting | Refuse to encode as |
|---|---|---|
| “Which screen gives you the id?” | Provenance | order_id with no description |
| “What do you never click from this queue?” | Negative trigger | A second write with the same stem |
| “What values does that dropdown actually send?” | Enum source | Marketing labels (Paused, On Hold) |
| “What happens if notes is blank?” | Optional vs required | A required string the handler defaults |
| “Who else can do this, and with which tool?” | Sibling / overlap | update_order “just in case” |
A schema written from the product glossary will require customer_name because sales likes names. The handler keys off order_id. The model will stuff a display name into an id field and you will 404 in a loop.
Do not write JSON until the card names a sibling tool. If there is no sibling, you probably hid a second side effect inside this one.
How do I name tools so they do not overlap?
Overlapping tools are two (or more) names whose descriptions could both match the same user utterance. The model then picks by vibe. That is not a reasoning failure. That is a catalog you authored badly.
Overlap test — run it out loud:
- Write five utterances the job actually hears (
"hold this order","stop the shipment","cancel it","refund the card","what's the status"). - For each utterance, list every tool whose current description could apply.
- Count greater than one is a fail. Zero on a write you expected is also a fail.
- Fix by renaming, adding a negative trigger, splitting the noun, or deleting the vaguer tool.
- Re-run the five utterances. Do not add a sixth tool until these five are unique.
| Utterance | Overlapping pair | Fix |
|---|---|---|
| “Hold this order” | update_order + hold_order | Delete general update_order on this job, or negative-trigger it |
| “Cancel it” | hold_order + cancel_order | Description: hold is reversible; cancel is not |
| “Email the customer” | send_email + send_ticket_reply | Noun in the name; send_email denied for this principal |
| “Find the order” | search_orders + get_order + lookup_order | Keep search (list) and get (one id). Delete lookup_ |
| “Fix the address” | update_order + update_shipping_address | Specific write tool; general update gone |
Naming rules that survive production:
verb_nounthe handler implements (hold_order, notorders, notcrm).- Asymmetric stems for read vs write:
search_vshold_. Same stem on both is how writes happen when you meant a lookup. - Do not ship
update_recordnext toupdate_order. The vaguer one wins under ambiguity. - Do not ship synonyms (
lookup_order/get_order/fetch_order). Pick one. - OWASP LLM06:2025 Excessive Agency is the risk name for too much functionality in one agent. Overlap is how that risk shows up in tool choice.
Fix menu when the overlap test fails — pick one, do not stack all four:
| Fix | Use when | Cost if you pick wrong |
|---|---|---|
| Delete the vaguer tool | update_order next to a specific write | You keep a confused deputy “for emergencies” |
| Rename the noun | Two legitimate writes, same verb | The model still matches on the old stem in memory of the prompt |
| Negative trigger by tool name | Both writes must exist | A soft “do not misuse” with no sibling named |
| Deny in the policy catalog | Irreversible twin on this principal | Description-only deny the model can ignore |
If two calls share a name but not a blast radius, they are two tools. If two names share a blast radius, they are one tool with a confused alias.
How do I write descriptions that constrain the call?
A constraining description answers four questions the model asks every turn. If any answer is missing, the description is decoration.
Template — fill every slot:
Call hold_order when the operator confirmed a reversible hold on an existing order
and search_orders already returned order_id this turn.
Do not call to cancel, refund, change address, or create an order.
Side effect: the order stops fulfillment until release_hold.
Sibling: cancel_order for irreversible cancel; refund_order for card movement.
| Slot | Job | Weak fill |
|---|---|---|
| When | Trigger | “Use this for orders” |
| When-not | Negative trigger | Omitted; model improvises |
| Side effect | What changes if this succeeds | “Updates a record” |
| Sibling | The overlapping tool you already killed on paper | “See other tools” |
Anthropic wants this in several sentences on the tool itself, not only in a system prompt the model can ignore when the tool list is long. OpenAI’s field table is the same idea in fewer words: when and how. Gemini wants you to be specific and to give examples. Three vendors, one job: the description is the tool-choice UX.
Checklist before you save the top-level string:
- Trigger is an observable state (confirmed hold, id already in context), not a mood (“when it seems right”)
- Negative trigger names the sibling tools by name
- Side effect is a business event, not “returns JSON”
- Id provenance is in the property text and hinted here (
from search_orders) - You did not paste the property list into the tool description — that is what
propertiesis for
A description that only restates the name ("Holds an order.") is how hold_order and cancel_order collapse into one guess.
How do I set required fields that match the handler?
required is a contract with two readers. If they disagree, you pay in invented args or 500s.
Procedure:
- Open the handler. List every key it reads. Mark each
must exist,defaulted, orignored. must exist→required, non-null type.defaulted→ not required on a non-strict provider; on OpenAI strict, keep the key, allownull, put it inrequired. OpenAI’s function calling strict rules: every object setsadditionalProperties: false, and every property is listed inrequired.ignored→ delete from the schema. A required field the handler ignores is an invitation to invent filler.- CI: diff
requiredagainst the handler destructure. A silent add from the exporter fails the build.
| Schema says | Handler does | What you will see |
|---|---|---|
note required string | Treats missing note as fine | Invented “Customer requested hold.” |
note omitted from required | Throws if note absent | Intermittent 500s, retry storm |
strict: true, note not required | Never reached | Provider 400 on the request |
note required, type ["string","null"] | Null means no note | Portable optional |
Worked hold tool (strict-safe: every property listed, optionals nullable):
{
"type": "object",
"additionalProperties": false,
"properties": {
"order_id": {
"type": "string",
"description": "Order UUID returned by search_orders this turn. Never invent an id. Never pass a confirmation number, email, or display name."
},
"hold_reason": {
"type": "string",
"enum": ["payment_review", "address_fix", "fraud_check", "customer_request"],
"description": "Exact hold_reason the orders API accepts today. Do not paraphrase. Do not send Complete or other."
},
"note": {
"type": ["string", "null"],
"description": "Optional internal note, max 280 characters. Null if the operator did not supply a note. Do not invent a note to satisfy required."
}
},
"required": ["order_id", "hold_reason", "note"]
}
Dishonest required is the quietest way to make an agent keep calling the same tool. The complementary diagnosis post covers the retry shape. Here the fix is mechanical: the array matches the handler, or you change the handler.
Write the error the model should see. An opaque 500 teaches a grind. A field-level string teaches one correction.
| Failure | Return | Do not return |
|---|---|---|
Missing order_id | order_id is required; pass the UUID from search_orders | Bad request |
| Invented id | order_id not found in this tenant; do not invent; call search_orders | 404 with no body |
| Enum miss | hold_reason 'fraud' is not allowed; use one of: payment_review, address_fix, fraud_check, customer_request | Stack trace |
| Wrong tool | status is not a field on update_order; to stop fulfillment call hold_order | Retry the same tool |
| Optional note invented | Do not 400 — if you required a string, that is your bug | Shame the model in the prompt |
Error-writing checklist:
- Field name in the string
- Rejected value echoed
- Allowed set or format named
- Sibling tool named when the miss was overlap, not type
- Auth and policy failures escalate; they are not schema retries
When should a field be an enum instead of a string?
Whenever the handler or the upstream API accepts a finite set. A free string on that field is a scheduled bug. JSON Schema enum exists to make the other values unrepresentable.
| Field class | Free string the model invents | Enum you write |
|---|---|---|
| Hold reason | fraud, waiting on payment, other | Values the orders API documents today |
| Channel | text, iMessage, sms (trailing space) | The two handlers you actually run |
| Cancel code | CustomerChangedMind | Snake_case the switch uses |
| Environment | prod, production, live | staging / prod if those are the two |
Enum writing rules:
- Copy from the handler
switchor the live API, not from a style guide. - Do not add synonyms (
completeanddone) unless both hit the same code path on purpose. - Put
"type": "string"next toenum. Gemini’s declarations are an OpenAPI subset; an enum without an explicit string type has 400’d real requests. - Version the enum when upstream adds a value. A frozen enum plus a new legitimate reason is a retry loop.
- Tool errors list the allowed set.
invalid hold_reason 'fraud'; allowed: payment_review, address_fix, fraud_check, customer_requestteaches one correction.400 Bad Requestteaches a grind.
Do not “helpfully” describe the allowed values only in prose ("one of payment review, address fix, …"). The sampler does not treat that sentence as a grammar. enum does.
If the set is not actually finite — free-text comments, user-authored subjects — do not fake an enum of five guesses. Constrain length and provenance in the description, and validate in the handler.
Version the enum like an API. A schema copy is a snapshot.
- Store
enum_versionnext to the tool (date or upstream hash). - Contract-test against the live or stubbed API enum endpoint in CI.
- When the test fails, do not hotfix to a free string. Add the value, add a golden row, then deploy.
- Alert on repeated
invalid_enumfor a value that is now legal upstream — that is drift, not a dumb model. - Keep a changelog line:
hold_reasonaddedcustomer_requeston 2026-06-12 because the orders API did.
An enum two releases behind the API is how a correct model enters a retry loop. Widening to a string to “unblock Friday” is how you re-open invented reasons on Monday.
How do I write per-property descriptions that constrain?
The property name is a hint. The description is the constraint. Names lie (amount, date, user, status).
Four sentences worth of constraint, even if you compress them into two:
| Constraint | Example on order_id | Example on note |
|---|---|---|
| Unit / format | UUID string, not integer | Max 280 characters |
| Provenance | From search_orders this turn | From the operator, or null |
| Forbidden substitute | Confirmation number, email, “the last one” | Invented summary of the ticket |
Example that matches type | "3f2c…" as a string | null when empty, not "" unless the handler wants empty string |
Property checklist — fail the review if any box is open on a write-path field:
- Unit named (
minutes,cents,ISO-8601 datetime,UUID) - Provenance named (
from search_orders, not “the id”) - Negative named (
Never invent,Do not pass a display name) - Example matches
type(integer example for integer fields; do not show"30"for an integer) - Null meaning named if the type includes
null
Empty property descriptions force the model to guess units. That is the complementary post’s production-bug table. The writing fix is this checklist, enforced in CI: a generator that emits a property without a description string fails the build.
Put closed sets in enum, not in the property text. Use the text for units and provenance. If you find yourself listing allowed values in a sentence, you still have not written the enum.
How do I close objects and skip keywords the sampler will ignore?
Close every object. Do not bet writes on keywords the provider accepts and the sampler does not enforce.
JSON Schema objects allow extra keys by default. Models use that slack. OpenAI strict rejects the request if any object in parameters omits additionalProperties: false. Nested objects count. Arrays of objects count. MCP’s tools spec recommends { "type": "object", "additionalProperties": false } for tools with no parameters so the only legal arguments object is {} — see the MCP tools page.
| Object | Closed? | Leak if you forget |
|---|---|---|
Root parameters / input_schema / inputSchema | Must be | Mystery top-level keys |
Nested address | Must be | address.notes the handler ignores |
items object in an array | Must be | Per-row junk |
| No-arg tool | Closed empty object | { "ok": true } anyway |
Keywords to describe in text and enforce in the handler — do not treat them as generation guarantees:
| Keyword | Typical fate | Write this instead |
|---|---|---|
pattern | Accepted or ignored; not a sure grammar | Description of the shape + handler reject |
minimum / maximum | Same | "1–100 inclusive" in the description + handler bounds |
format: date-time | Hint | ISO-8601 example + handler |
$ref / deep $defs | Often rejected in strict | Flatten for the provider export |
Portable rule: describe bounds, enforce bounds in the handler, return a field-level error. The schema is the model’s UX. The handler is the second reader. Closing the object is still not authorization — a well-typed invented UUID is still an invented UUID.
How do I evaluate this in production?
You evaluate schemas with a golden set that can fail tool choice and argument shape separately. A single “task passed” number will hide a schema that still burns retries. That is the same independence rule as evaluators before agents: the judge is not the worker, and the metric is not a vibe.
Minimum columns:
| Column | Example |
|---|---|
utterance | “Hold order 3f2c — payment review. No note.” |
expected_tool | hold_order |
forbidden_tools | cancel_order, refund_order, update_order |
expected_args | { "order_id": "3f2c…", "hold_reason": "payment_review", "note": null } |
almost_right | hold_reason: "fraud" (not in enum) |
Score separately, every schema change:
- Tool-choice accuracy — right name. Overlap traps live here.
- Schema-valid args — types, required keys, enum membership.
- Argument exact-match after normalizing nulls and UUID case.
- Forbidden-tool rate — the overlap sheet, scored, not hoped.
Eval hygiene:
- Compiler errors (provider rejected the schema) are a separate bucket from model errors
- Each write tool has one almost-right case (
fraudvsfraud_check, confirmation number vs UUID) - Each overlapping pair from the overlap test has a case
- You re-run after enum or exporter changes, not only after prompt edits
- Online: alert on repeated
invalid_enum/ missing-required tool errors for the same tool name
I will not mint a studio-wide “schema quality score.” Yours is: share of golden rows that match expected tool and expected args, plus the forbidden-tool rate. If you cannot name the last time argument exact-match moved, you are flying on anecdotes.
A schema change that lifts exact-match while task pass rate stays flat still shipped. Fewer retries. Fewer weird writes. Do not bury that lift inside one success percentage you cannot source.
What guardrails do I need?
A schema the model obeys is necessary and not sufficient. Guardrails sit around the call.
| Guardrail | What it stops | What the schema cannot do instead |
|---|---|---|
| Overlap-free catalog | Wrong write under a vague utterance | A longer description on a god-tool |
Honest required + enums | Invented notes and invented stages | A prompt that says “be careful with ids” |
| Closed objects | Mystery keys | Hoping the handler ignores them |
| Handler validation | Well-typed invented UUIDs, amounts over a cap | maximum the sampler may not enforce |
| Policy allow / deny / pending-approval | Unauthorized side effect | additionalProperties: false as authz |
| Runtime-minted idempotency keys | Double send on retry | Letting the model invent the key |
Minimum this week on any write path:
- Overlap test signed off. No
update_*twins. - Zero empty property descriptions on tools that send email or move money.
- Enums copied from the live API, with a contract test so they cannot freeze.
- Handler errors that name the field and the allowed set.
- Golden set with overlap traps before the tool is allowlisted in prod.
additionalProperties: false is grammar. A grammatically perfect refund_order can still be a policy violation. Gate the side effect in the harness. The schema’s job is to stop illegal shapes, not illegal decisions.
When is a workflow enough instead of an agent?
A workflow is enough when the next tool is determined by the last result with no branching judgment. An agent is for jobs where tool choice is the product. Schema quality still matters on a workflow node. Overlap does not, because there is one write node.
| Signal | Ship a workflow | Keep an agent |
|---|---|---|
| Path | Lookup → hold → notify, always | Operator utterance could mean hold, cancel, or refund |
| Tool choice | One write tool on the node | Several writes; the model must pick |
| Failure | Retry the node with a clock | Fingerprint, cap, escalate — not a schema essay |
| Schema job | Args on that one node | Args and which tool |
If you are writing twenty tools so the model can “decide,” and the operator’s real path is three steps with a human on the irreversible one, you are paying tool-choice error for a flowchart. Write the three-node workflow. Put a strict schema on the write node. Keep the overlap test in your pocket for the day you actually need an agent.
A good schema on a workflow write node still needs required fields, enums, and constraining descriptions. You just do not also need a catalog of near-synonyms.
Decision procedure — stop at the first yes:
- Can you draw the path on a whiteboard with named nodes and no “the model picks”? → workflow.
- Is the only uncertainty which row to load, not which write? → workflow plus a search node.
- Does the operator utterance branch across hold / cancel / refund / address with real judgment? → agent, overlap test mandatory.
- Is the irreversible step always a human? → workflow with a wait node, not an agent that “knows when to ask.”
- Are you adding tools so the demo looks autonomous? → delete them; you are buying overlap.
Agents pay a tool-choice tax. Do not pay it for a flowchart you already run at 2 a.m. without a model.
Failure mode: two tools that can both “fix the order”
What broke: The catalog shipped update_order (open payload object) next to hold_order (tight enum). Operators said “stop this order.” The model picked update_order, invented status: "paused", and the API 400’d. The agent retried update_order because the error did not name hold_order.
What it cost: A hold that never landed, a fulfillment window that closed, and a human who “fixed” it by widening status to a free string — which reopened invented statuses on the general update tool.
Fix path:
- Delete or deny
update_orderon this job. Specific writes only. - Negative-trigger
hold_orderagainstcancel_orderandrefund_orderby name. - Return
unknown status 'paused'; this job uses hold_order with hold_reason one of: …so the model can switch tools once. - Add the utterance to the golden set as an overlap trap.
- Refuse the “just make status a string” hotfix on the write path.
Second failure we keep seeing: someone adds note to required as a string because “the CRM form has a notes box.” The model invents a paragraph every call. Support starts treating those notes as operator-authored. Make note required-plus-null, describe “do not invent,” and stop scoring “the agent always fills notes” as a win.
Third failure: nested objects with empty inner schemas. shipping: { "type": "object" } with no properties is an omnibus one level down. The model invents shipping.notes, shipping.priority, shipping.leave_at_door. Close the nested object, list the keys the handler reads, enum the service level, and describe the provenance of address_id the same way you describe order_id. Root-level discipline that stops at the first nested bag is not discipline.
One-week writing order
Skip Pydantic/Zod exporter work, MCP catalog dumps, and a 200-case eval harness. Those are later. This week you write the schemas the write path actually uses.
| Day | Do | Do not |
|---|---|---|
| 1 | Side-effect cards + overlap test on current names | Add tools “for completeness” |
| 2 | Rewrite top-level descriptions (when / when-not / sibling) | Lengthen the system prompt instead |
| 3 | Align required with handlers; nullable optionals under strict | Require vanity fields the handler ignores |
| 4 | Enums from live APIs; property constraint checklist; close objects | pattern as a substitute for handler bounds |
| 5 | Field-level errors + 20–40 golden rows including overlap traps | Declare victory on one demo transcript |
Week-one skip list:
- Generating schemas from rich models without a description lint
- Loading every MCP server tool into the active set
- Omnibus
action+payload“for flexibility” - Synonym tools (
lookup_/get_/fetch_) - Free-string status fields on money or fulfillment writes
In a Spurlock five-day agentic pilot: inventory and kill overlapping writes, rewrite required / enum / descriptions on the remaining write path, and return a golden slice that scores tool choice and argument exact-match. Zero empty property descriptions on tools that send email or move money.
FAQ
How do I write good tool schemas for AI agents?
Write side-effect cards and run an overlap test before JSON, then project each card into a verb_noun tool with a when/when-not description, a required array that matches the handler, enums copied from the live API, and a constraining description on every property. Close objects with additionalProperties: false. The schema is agent-facing UX. A longer prompt will not repair two tools that both mean “fix the order.”
How do I measure whether do I write good tool schemas for AI agents is working?
Score a golden set for tool-choice accuracy, schema-valid args, argument exact-match, and forbidden-tool rate — separately, not as one task-pass percentage. Include overlap traps and almost-right enums. Online, watch repeated validation errors on the same write tool. If exact-match never moves after a schema edit, you did not measure the schema; you measured the demo.
What usually fails first when teams try this?
Overlap, then dishonest required. Two writes that can match the same utterance, or a required string the handler does not need, will show up before anyone notices a missing additionalProperties. Third is a free-string field that should have been an enum, which turns into a retry loop the first time upstream rejects the paraphrase.
How long does this take to show results?
A week is enough to rewrite the write-path schemas for one job and see argument exact-match move on a 20–40 row golden set. You will see production results the first time a hold uses a real order_id and a legal hold_reason instead of a paraphrased status on a general update tool. Broader catalogs are a calendar. I will not invent a lift percentage.
What should I skip if I only have a week?
Skip exporter rewrites, loading every MCP tool, and synonym aliases. Do not skip the overlap test, handler-true required, enums on closed sets, property descriptions that name provenance, closed objects on writes, and golden cases for the scary utterances. Unstaffed irreversible tools stay denied. A prettier demo schema is not a week.
When is this not worth doing yet?
If the job is a fixed lookup-then-write flowchart, ship a workflow with one strict schema on the write node instead of an agent catalog. If you have no side-effecting tools — no send, no hold, no refund — you do not have a schema-quality problem yet; you have a chatbot. If nobody will run the overlap test or own the enum when the API adds a value, do not allowlist production writes.
CTA
Want write-path schemas that stop overlapping tools and invented args on a real job in five days?
What questions does this article answer?
- How do I write good tool schemas for AI agents?
- Write side-effect cards and run an overlap test before JSON, then project each card into a `verb_noun` tool with a when/when-not description, a `required` array that matches the handler, enums copied from the live API, and a constraining description on every property. Close objects with `additionalProperties: false`. The schema is agent-facing UX. A longer prompt will not repair two tools that both mean “fix the order.”
- How do I measure whether do I write good tool schemas for AI agents is working?
- Score a golden set for tool-choice accuracy, schema-valid args, argument exact-match, and forbidden-tool rate — separately, not as one task-pass percentage. Include overlap traps and almost-right enums. Online, watch repeated validation errors on the same write tool. If exact-match never moves after a schema edit, you did not measure the schema; you measured the demo.
- What usually fails first when teams try this?
- Overlap, then dishonest `required`. Two writes that can match the same utterance, or a required string the handler does not need, will show up before anyone notices a missing `additionalProperties`. Third is a free-string field that should have been an enum, which turns into a retry loop the first time upstream rejects the paraphrase.
- How long does this take to show results?
- A week is enough to rewrite the write-path schemas for one job and see argument exact-match move on a 20–40 row golden set. You will see production results the first time a hold uses a real `order_id` and a legal `hold_reason` instead of a paraphrased status on a general update tool. Broader catalogs are a calendar. I will not invent a lift percentage.
- What should I skip if I only have a week?
- Skip exporter rewrites, loading every MCP tool, and synonym aliases. Do not skip the overlap test, handler-true `required`, enums on closed sets, property descriptions that name provenance, closed objects on writes, and golden cases for the scary utterances. Unstaffed irreversible tools stay denied. A prettier demo schema is not a week.
- When is this not worth doing yet?
- If the job is a fixed lookup-then-write flowchart, ship a workflow with one strict schema on the write node instead of an agent catalog. If you have no side-effecting tools — no send, no hold, no refund — you do not have a schema-quality problem yet; you have a chatbot. If nobody will run the overlap test or own the enum when the API adds a value, do not allowlist production writes.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents When should the agent escalate instead of retrying
Escalate on auth, policy, ambiguous intent, repeated same-tool fail, and money movement. Retry only transient, idempotent tool errors with a hard bound.
AI Agents Why is my agent 10× more expensive than the chatbot demo
Agents cost more than the chatbot demo because each tool turn re-bills growing context, schemas, and retries. 10× is a complaint to diagnose, not a statistic.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.