Spurlock Studios
Contact
Share LinkedIn X
Nested brass frames. Thesis: WRITE GOOD TOOL SCHEMAS AI.

You write good tool schemas for AI agents by treating JSON Schema as agent-facing UX: one tool per side effect, an honest required array, enums for every closed set, and property descriptions that name units, provenance, and forbidden substitutes. Vague names plus empty descriptions are how the model invents arguments. Overlapping tools are how it picks the wrong write. This is a writing procedure, not a longer system prompt.

This spoke sits under the Agentic Systems Operating Manual. It is the how-to. The diagnosis of vague schemas, strict-mode grammar, and the omnibus do_anything tool lives in tool schemas agents follow. Score the result the way you build the evaluator before the agent: a golden set that can fail tool choice and argument shape separately.

The short answer

  • Inventory side effects first. JSON is last.
  • One verb_noun per blast radius. If two tools can satisfy the same user sentence, they overlap — rename, add a negative trigger, or delete one.
  • required matches the handler. Under OpenAI strict mode, optional-in-product is required-plus-null.
  • Finite values are enum, copied from the switch, not from the glossary.
  • Every property description constrains: unit, where the id came from, what never to pass, one example that matches type.
You are writingDone looks likeNot done
Namehold_order, not ordersTwo stems that both mean “fix the order”
Top-level descriptionWhen / when-not / side effect / sibling toolA restatement of the function name
requiredKeys the handler cannot run without“Nice to have” fields the model will invent
Closed setenum equal to the APIA comma list inside the description
Property textUnit + provenance + never-inventEmpty string, or the name repeated

How do I write good tool schemas for AI agents?

You write them in a fixed order so the model never sees a tool whose handler and schema disagree. Do not start in the JSON editor.

Writing order:

  1. List every side effect this job is allowed to perform (read, write_reversible, write_irreversible, exfil_risk).
  2. Run the overlap test on the names you want to ship (next section). Zero overlaps, or you stop.
  3. Name each tool verb_noun the handler actually implements.
  4. Write the top-level description as trigger language — OpenAI’s function calling field table calls description “details on when and how to use the function.”
  5. Open the handler. Mark every argument it cannot default. Those keys are required.
  6. For each finite set, copy the handler switch / upstream enum into enum. Gemini’s function calling guide says to put closed sets in enum instead of only describing them.
  7. Write a constraining description on every property. Anthropic’s define tools page calls detailed descriptions “by far the most important factor in tool performance” and wants several sentences on the tool plus a description on each parameter.
  8. Close every object with additionalProperties: false. JSON Schema treats extra keys as valid unless you say otherwise.
  9. Write the handler error the model should see on a bad value: field name, rejected value, allowed set.
  10. Add golden cases: one happy path, one almost-right arg, one overlap trap.
StepArtifactReader
1–2Side-effect table + overlap sheetYou
3–4name + top-level descriptionModel (tool choice)
5–8required, enum, property text, closed objectsModel (argument shape)
9–10Error strings + golden setHandler + evaluator

If you skip to step 7 because “the types are obvious,” you will ship a legal schema the model still mis-calls. Types without triggers are a dictionary, not a tool.

What do I write down before any JSON Schema?

You write a side-effect card per tool. The card is the source of truth. The schema is a projection of the card into JSON Schema.

Card fieldExampleWhy it exists
namehold_orderTool-choice string
Side-effect classwrite_reversiblePolicy gate, not the model
Blast radiusOne order, reversible by release_holdOverlap test
TriggerOperator confirmed a hold; order_id already from search_ordersTop-level description
Negative triggerDo not use to cancel, refund, or change addressSibling tools
Handler cannot-run-withoutorder_id, hold_reasonrequired
Closed setshold_reason ∈ four API valuesenum
Id provenancesearch_orders hit this turnProperty description
Forbidden substitutesConfirmation number, email, “the last order”Property description

Card procedure:

  1. Sit with the operator who already does this job by hand. Ask what they click, what they refuse, and which id they copy from which screen.
  2. Fill the card in their words. If they cannot name a negative trigger, the tool is still an omnibus in disguise.
  3. Only then open the schema file.
  4. Diff the card against the handler signature. The card loses. The handler wins.

Interview questions that actually fill the card:

AskYou are extractingRefuse to encode as
“Which screen gives you the id?”Provenanceorder_id with no description
“What do you never click from this queue?”Negative triggerA second write with the same stem
“What values does that dropdown actually send?”Enum sourceMarketing labels (Paused, On Hold)
“What happens if notes is blank?”Optional vs requiredA required string the handler defaults
“Who else can do this, and with which tool?”Sibling / overlapupdate_order “just in case”

A schema written from the product glossary will require customer_name because sales likes names. The handler keys off order_id. The model will stuff a display name into an id field and you will 404 in a loop.

Do not write JSON until the card names a sibling tool. If there is no sibling, you probably hid a second side effect inside this one.

How do I name tools so they do not overlap?

Overlapping tools are two (or more) names whose descriptions could both match the same user utterance. The model then picks by vibe. That is not a reasoning failure. That is a catalog you authored badly.

Overlap test — run it out loud:

  1. Write five utterances the job actually hears ("hold this order", "stop the shipment", "cancel it", "refund the card", "what's the status").
  2. For each utterance, list every tool whose current description could apply.
  3. Count greater than one is a fail. Zero on a write you expected is also a fail.
  4. Fix by renaming, adding a negative trigger, splitting the noun, or deleting the vaguer tool.
  5. Re-run the five utterances. Do not add a sixth tool until these five are unique.
UtteranceOverlapping pairFix
“Hold this order”update_order + hold_orderDelete general update_order on this job, or negative-trigger it
“Cancel it”hold_order + cancel_orderDescription: hold is reversible; cancel is not
“Email the customer”send_email + send_ticket_replyNoun in the name; send_email denied for this principal
“Find the order”search_orders + get_order + lookup_orderKeep search (list) and get (one id). Delete lookup_
“Fix the address”update_order + update_shipping_addressSpecific write tool; general update gone

Naming rules that survive production:

  • verb_noun the handler implements (hold_order, not orders, not crm).
  • Asymmetric stems for read vs write: search_ vs hold_. Same stem on both is how writes happen when you meant a lookup.
  • Do not ship update_record next to update_order. The vaguer one wins under ambiguity.
  • Do not ship synonyms (lookup_order / get_order / fetch_order). Pick one.
  • OWASP LLM06:2025 Excessive Agency is the risk name for too much functionality in one agent. Overlap is how that risk shows up in tool choice.

Fix menu when the overlap test fails — pick one, do not stack all four:

FixUse whenCost if you pick wrong
Delete the vaguer toolupdate_order next to a specific writeYou keep a confused deputy “for emergencies”
Rename the nounTwo legitimate writes, same verbThe model still matches on the old stem in memory of the prompt
Negative trigger by tool nameBoth writes must existA soft “do not misuse” with no sibling named
Deny in the policy catalogIrreversible twin on this principalDescription-only deny the model can ignore

If two calls share a name but not a blast radius, they are two tools. If two names share a blast radius, they are one tool with a confused alias.

How do I write descriptions that constrain the call?

A constraining description answers four questions the model asks every turn. If any answer is missing, the description is decoration.

Template — fill every slot:

Call hold_order when the operator confirmed a reversible hold on an existing order
and search_orders already returned order_id this turn.
Do not call to cancel, refund, change address, or create an order.
Side effect: the order stops fulfillment until release_hold.
Sibling: cancel_order for irreversible cancel; refund_order for card movement.
SlotJobWeak fill
WhenTrigger“Use this for orders”
When-notNegative triggerOmitted; model improvises
Side effectWhat changes if this succeeds“Updates a record”
SiblingThe overlapping tool you already killed on paper“See other tools”

Anthropic wants this in several sentences on the tool itself, not only in a system prompt the model can ignore when the tool list is long. OpenAI’s field table is the same idea in fewer words: when and how. Gemini wants you to be specific and to give examples. Three vendors, one job: the description is the tool-choice UX.

Checklist before you save the top-level string:

  • Trigger is an observable state (confirmed hold, id already in context), not a mood (“when it seems right”)
  • Negative trigger names the sibling tools by name
  • Side effect is a business event, not “returns JSON”
  • Id provenance is in the property text and hinted here (from search_orders)
  • You did not paste the property list into the tool description — that is what properties is for

A description that only restates the name ("Holds an order.") is how hold_order and cancel_order collapse into one guess.

How do I set required fields that match the handler?

required is a contract with two readers. If they disagree, you pay in invented args or 500s.

Procedure:

  1. Open the handler. List every key it reads. Mark each must exist, defaulted, or ignored.
  2. must exist → required, non-null type.
  3. defaulted → not required on a non-strict provider; on OpenAI strict, keep the key, allow null, put it in required. OpenAI’s function calling strict rules: every object sets additionalProperties: false, and every property is listed in required.
  4. ignored → delete from the schema. A required field the handler ignores is an invitation to invent filler.
  5. CI: diff required against the handler destructure. A silent add from the exporter fails the build.
Schema saysHandler doesWhat you will see
note required stringTreats missing note as fineInvented “Customer requested hold.”
note omitted from requiredThrows if note absentIntermittent 500s, retry storm
strict: true, note not requiredNever reachedProvider 400 on the request
note required, type ["string","null"]Null means no notePortable optional

Worked hold tool (strict-safe: every property listed, optionals nullable):

{
  "type": "object",
  "additionalProperties": false,
  "properties": {
    "order_id": {
      "type": "string",
      "description": "Order UUID returned by search_orders this turn. Never invent an id. Never pass a confirmation number, email, or display name."
    },
    "hold_reason": {
      "type": "string",
      "enum": ["payment_review", "address_fix", "fraud_check", "customer_request"],
      "description": "Exact hold_reason the orders API accepts today. Do not paraphrase. Do not send Complete or other."
    },
    "note": {
      "type": ["string", "null"],
      "description": "Optional internal note, max 280 characters. Null if the operator did not supply a note. Do not invent a note to satisfy required."
    }
  },
  "required": ["order_id", "hold_reason", "note"]
}

Dishonest required is the quietest way to make an agent keep calling the same tool. The complementary diagnosis post covers the retry shape. Here the fix is mechanical: the array matches the handler, or you change the handler.

Write the error the model should see. An opaque 500 teaches a grind. A field-level string teaches one correction.

FailureReturnDo not return
Missing order_idorder_id is required; pass the UUID from search_ordersBad request
Invented idorder_id not found in this tenant; do not invent; call search_orders404 with no body
Enum misshold_reason 'fraud' is not allowed; use one of: payment_review, address_fix, fraud_check, customer_requestStack trace
Wrong toolstatus is not a field on update_order; to stop fulfillment call hold_orderRetry the same tool
Optional note inventedDo not 400 — if you required a string, that is your bugShame the model in the prompt

Error-writing checklist:

  • Field name in the string
  • Rejected value echoed
  • Allowed set or format named
  • Sibling tool named when the miss was overlap, not type
  • Auth and policy failures escalate; they are not schema retries

When should a field be an enum instead of a string?

Whenever the handler or the upstream API accepts a finite set. A free string on that field is a scheduled bug. JSON Schema enum exists to make the other values unrepresentable.

Field classFree string the model inventsEnum you write
Hold reasonfraud, waiting on payment, otherValues the orders API documents today
Channeltext, iMessage, sms (trailing space)The two handlers you actually run
Cancel codeCustomerChangedMindSnake_case the switch uses
Environmentprod, production, livestaging / prod if those are the two

Enum writing rules:

  1. Copy from the handler switch or the live API, not from a style guide.
  2. Do not add synonyms (complete and done) unless both hit the same code path on purpose.
  3. Put "type": "string" next to enum. Gemini’s declarations are an OpenAPI subset; an enum without an explicit string type has 400’d real requests.
  4. Version the enum when upstream adds a value. A frozen enum plus a new legitimate reason is a retry loop.
  5. Tool errors list the allowed set. invalid hold_reason 'fraud'; allowed: payment_review, address_fix, fraud_check, customer_request teaches one correction. 400 Bad Request teaches a grind.

Do not “helpfully” describe the allowed values only in prose ("one of payment review, address fix, …"). The sampler does not treat that sentence as a grammar. enum does.

If the set is not actually finite — free-text comments, user-authored subjects — do not fake an enum of five guesses. Constrain length and provenance in the description, and validate in the handler.

Version the enum like an API. A schema copy is a snapshot.

  1. Store enum_version next to the tool (date or upstream hash).
  2. Contract-test against the live or stubbed API enum endpoint in CI.
  3. When the test fails, do not hotfix to a free string. Add the value, add a golden row, then deploy.
  4. Alert on repeated invalid_enum for a value that is now legal upstream — that is drift, not a dumb model.
  5. Keep a changelog line: hold_reason added customer_request on 2026-06-12 because the orders API did.

An enum two releases behind the API is how a correct model enters a retry loop. Widening to a string to “unblock Friday” is how you re-open invented reasons on Monday.

How do I write per-property descriptions that constrain?

The property name is a hint. The description is the constraint. Names lie (amount, date, user, status).

Four sentences worth of constraint, even if you compress them into two:

ConstraintExample on order_idExample on note
Unit / formatUUID string, not integerMax 280 characters
ProvenanceFrom search_orders this turnFrom the operator, or null
Forbidden substituteConfirmation number, email, “the last one”Invented summary of the ticket
Example that matches type"3f2c…" as a stringnull when empty, not "" unless the handler wants empty string

Property checklist — fail the review if any box is open on a write-path field:

  • Unit named (minutes, cents, ISO-8601 datetime, UUID)
  • Provenance named (from search_orders, not “the id”)
  • Negative named (Never invent, Do not pass a display name)
  • Example matches type (integer example for integer fields; do not show "30" for an integer)
  • Null meaning named if the type includes null

Empty property descriptions force the model to guess units. That is the complementary post’s production-bug table. The writing fix is this checklist, enforced in CI: a generator that emits a property without a description string fails the build.

Put closed sets in enum, not in the property text. Use the text for units and provenance. If you find yourself listing allowed values in a sentence, you still have not written the enum.

How do I close objects and skip keywords the sampler will ignore?

Close every object. Do not bet writes on keywords the provider accepts and the sampler does not enforce.

JSON Schema objects allow extra keys by default. Models use that slack. OpenAI strict rejects the request if any object in parameters omits additionalProperties: false. Nested objects count. Arrays of objects count. MCP’s tools spec recommends { "type": "object", "additionalProperties": false } for tools with no parameters so the only legal arguments object is {} — see the MCP tools page.

ObjectClosed?Leak if you forget
Root parameters / input_schema / inputSchemaMust beMystery top-level keys
Nested addressMust beaddress.notes the handler ignores
items object in an arrayMust bePer-row junk
No-arg toolClosed empty object{ "ok": true } anyway

Keywords to describe in text and enforce in the handler — do not treat them as generation guarantees:

KeywordTypical fateWrite this instead
patternAccepted or ignored; not a sure grammarDescription of the shape + handler reject
minimum / maximumSame"1–100 inclusive" in the description + handler bounds
format: date-timeHintISO-8601 example + handler
$ref / deep $defsOften rejected in strictFlatten for the provider export

Portable rule: describe bounds, enforce bounds in the handler, return a field-level error. The schema is the model’s UX. The handler is the second reader. Closing the object is still not authorization — a well-typed invented UUID is still an invented UUID.

How do I evaluate this in production?

You evaluate schemas with a golden set that can fail tool choice and argument shape separately. A single “task passed” number will hide a schema that still burns retries. That is the same independence rule as evaluators before agents: the judge is not the worker, and the metric is not a vibe.

Minimum columns:

ColumnExample
utterance“Hold order 3f2c — payment review. No note.”
expected_toolhold_order
forbidden_toolscancel_order, refund_order, update_order
expected_args{ "order_id": "3f2c…", "hold_reason": "payment_review", "note": null }
almost_righthold_reason: "fraud" (not in enum)

Score separately, every schema change:

  1. Tool-choice accuracy — right name. Overlap traps live here.
  2. Schema-valid args — types, required keys, enum membership.
  3. Argument exact-match after normalizing nulls and UUID case.
  4. Forbidden-tool rate — the overlap sheet, scored, not hoped.

Eval hygiene:

  • Compiler errors (provider rejected the schema) are a separate bucket from model errors
  • Each write tool has one almost-right case (fraud vs fraud_check, confirmation number vs UUID)
  • Each overlapping pair from the overlap test has a case
  • You re-run after enum or exporter changes, not only after prompt edits
  • Online: alert on repeated invalid_enum / missing-required tool errors for the same tool name

I will not mint a studio-wide “schema quality score.” Yours is: share of golden rows that match expected tool and expected args, plus the forbidden-tool rate. If you cannot name the last time argument exact-match moved, you are flying on anecdotes.

A schema change that lifts exact-match while task pass rate stays flat still shipped. Fewer retries. Fewer weird writes. Do not bury that lift inside one success percentage you cannot source.

What guardrails do I need?

A schema the model obeys is necessary and not sufficient. Guardrails sit around the call.

GuardrailWhat it stopsWhat the schema cannot do instead
Overlap-free catalogWrong write under a vague utteranceA longer description on a god-tool
Honest required + enumsInvented notes and invented stagesA prompt that says “be careful with ids”
Closed objectsMystery keysHoping the handler ignores them
Handler validationWell-typed invented UUIDs, amounts over a capmaximum the sampler may not enforce
Policy allow / deny / pending-approvalUnauthorized side effectadditionalProperties: false as authz
Runtime-minted idempotency keysDouble send on retryLetting the model invent the key

Minimum this week on any write path:

  1. Overlap test signed off. No update_* twins.
  2. Zero empty property descriptions on tools that send email or move money.
  3. Enums copied from the live API, with a contract test so they cannot freeze.
  4. Handler errors that name the field and the allowed set.
  5. Golden set with overlap traps before the tool is allowlisted in prod.

additionalProperties: false is grammar. A grammatically perfect refund_order can still be a policy violation. Gate the side effect in the harness. The schema’s job is to stop illegal shapes, not illegal decisions.

When is a workflow enough instead of an agent?

A workflow is enough when the next tool is determined by the last result with no branching judgment. An agent is for jobs where tool choice is the product. Schema quality still matters on a workflow node. Overlap does not, because there is one write node.

SignalShip a workflowKeep an agent
PathLookup → hold → notify, alwaysOperator utterance could mean hold, cancel, or refund
Tool choiceOne write tool on the nodeSeveral writes; the model must pick
FailureRetry the node with a clockFingerprint, cap, escalate — not a schema essay
Schema jobArgs on that one nodeArgs and which tool

If you are writing twenty tools so the model can “decide,” and the operator’s real path is three steps with a human on the irreversible one, you are paying tool-choice error for a flowchart. Write the three-node workflow. Put a strict schema on the write node. Keep the overlap test in your pocket for the day you actually need an agent.

A good schema on a workflow write node still needs required fields, enums, and constraining descriptions. You just do not also need a catalog of near-synonyms.

Decision procedure — stop at the first yes:

  1. Can you draw the path on a whiteboard with named nodes and no “the model picks”? → workflow.
  2. Is the only uncertainty which row to load, not which write? → workflow plus a search node.
  3. Does the operator utterance branch across hold / cancel / refund / address with real judgment? → agent, overlap test mandatory.
  4. Is the irreversible step always a human? → workflow with a wait node, not an agent that “knows when to ask.”
  5. Are you adding tools so the demo looks autonomous? → delete them; you are buying overlap.

Agents pay a tool-choice tax. Do not pay it for a flowchart you already run at 2 a.m. without a model.

Failure mode: two tools that can both “fix the order”

What broke: The catalog shipped update_order (open payload object) next to hold_order (tight enum). Operators said “stop this order.” The model picked update_order, invented status: "paused", and the API 400’d. The agent retried update_order because the error did not name hold_order.

What it cost: A hold that never landed, a fulfillment window that closed, and a human who “fixed” it by widening status to a free string — which reopened invented statuses on the general update tool.

Fix path:

  1. Delete or deny update_order on this job. Specific writes only.
  2. Negative-trigger hold_order against cancel_order and refund_order by name.
  3. Return unknown status 'paused'; this job uses hold_order with hold_reason one of: … so the model can switch tools once.
  4. Add the utterance to the golden set as an overlap trap.
  5. Refuse the “just make status a string” hotfix on the write path.

Second failure we keep seeing: someone adds note to required as a string because “the CRM form has a notes box.” The model invents a paragraph every call. Support starts treating those notes as operator-authored. Make note required-plus-null, describe “do not invent,” and stop scoring “the agent always fills notes” as a win.

Third failure: nested objects with empty inner schemas. shipping: { "type": "object" } with no properties is an omnibus one level down. The model invents shipping.notes, shipping.priority, shipping.leave_at_door. Close the nested object, list the keys the handler reads, enum the service level, and describe the provenance of address_id the same way you describe order_id. Root-level discipline that stops at the first nested bag is not discipline.

One-week writing order

Skip Pydantic/Zod exporter work, MCP catalog dumps, and a 200-case eval harness. Those are later. This week you write the schemas the write path actually uses.

DayDoDo not
1Side-effect cards + overlap test on current namesAdd tools “for completeness”
2Rewrite top-level descriptions (when / when-not / sibling)Lengthen the system prompt instead
3Align required with handlers; nullable optionals under strictRequire vanity fields the handler ignores
4Enums from live APIs; property constraint checklist; close objectspattern as a substitute for handler bounds
5Field-level errors + 20–40 golden rows including overlap trapsDeclare victory on one demo transcript

Week-one skip list:

  • Generating schemas from rich models without a description lint
  • Loading every MCP server tool into the active set
  • Omnibus action + payload “for flexibility”
  • Synonym tools (lookup_ / get_ / fetch_)
  • Free-string status fields on money or fulfillment writes

In a Spurlock five-day agentic pilot: inventory and kill overlapping writes, rewrite required / enum / descriptions on the remaining write path, and return a golden slice that scores tool choice and argument exact-match. Zero empty property descriptions on tools that send email or move money.

FAQ

How do I write good tool schemas for AI agents?

Write side-effect cards and run an overlap test before JSON, then project each card into a verb_noun tool with a when/when-not description, a required array that matches the handler, enums copied from the live API, and a constraining description on every property. Close objects with additionalProperties: false. The schema is agent-facing UX. A longer prompt will not repair two tools that both mean “fix the order.”

How do I measure whether do I write good tool schemas for AI agents is working?

Score a golden set for tool-choice accuracy, schema-valid args, argument exact-match, and forbidden-tool rate — separately, not as one task-pass percentage. Include overlap traps and almost-right enums. Online, watch repeated validation errors on the same write tool. If exact-match never moves after a schema edit, you did not measure the schema; you measured the demo.

What usually fails first when teams try this?

Overlap, then dishonest required. Two writes that can match the same utterance, or a required string the handler does not need, will show up before anyone notices a missing additionalProperties. Third is a free-string field that should have been an enum, which turns into a retry loop the first time upstream rejects the paraphrase.

How long does this take to show results?

A week is enough to rewrite the write-path schemas for one job and see argument exact-match move on a 20–40 row golden set. You will see production results the first time a hold uses a real order_id and a legal hold_reason instead of a paraphrased status on a general update tool. Broader catalogs are a calendar. I will not invent a lift percentage.

What should I skip if I only have a week?

Skip exporter rewrites, loading every MCP tool, and synonym aliases. Do not skip the overlap test, handler-true required, enums on closed sets, property descriptions that name provenance, closed objects on writes, and golden cases for the scary utterances. Unstaffed irreversible tools stay denied. A prettier demo schema is not a week.

When is this not worth doing yet?

If the job is a fixed lookup-then-write flowchart, ship a workflow with one strict schema on the write node instead of an agent catalog. If you have no side-effecting tools — no send, no hold, no refund — you do not have a schema-quality problem yet; you have a chatbot. If nobody will run the overlap test or own the enum when the API adds a value, do not allowlist production writes.

CTA

Want write-path schemas that stop overlapping tools and invented args on a real job in five days?

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

How do I write good tool schemas for AI agents?
Write side-effect cards and run an overlap test before JSON, then project each card into a `verb_noun` tool with a when/when-not description, a `required` array that matches the handler, enums copied from the live API, and a constraining description on every property. Close objects with `additionalProperties: false`. The schema is agent-facing UX. A longer prompt will not repair two tools that both mean “fix the order.”
How do I measure whether do I write good tool schemas for AI agents is working?
Score a golden set for tool-choice accuracy, schema-valid args, argument exact-match, and forbidden-tool rate — separately, not as one task-pass percentage. Include overlap traps and almost-right enums. Online, watch repeated validation errors on the same write tool. If exact-match never moves after a schema edit, you did not measure the schema; you measured the demo.
What usually fails first when teams try this?
Overlap, then dishonest `required`. Two writes that can match the same utterance, or a required string the handler does not need, will show up before anyone notices a missing `additionalProperties`. Third is a free-string field that should have been an enum, which turns into a retry loop the first time upstream rejects the paraphrase.
How long does this take to show results?
A week is enough to rewrite the write-path schemas for one job and see argument exact-match move on a 20–40 row golden set. You will see production results the first time a hold uses a real `order_id` and a legal `hold_reason` instead of a paraphrased status on a general update tool. Broader catalogs are a calendar. I will not invent a lift percentage.
What should I skip if I only have a week?
Skip exporter rewrites, loading every MCP tool, and synonym aliases. Do not skip the overlap test, handler-true `required`, enums on closed sets, property descriptions that name provenance, closed objects on writes, and golden cases for the scary utterances. Unstaffed irreversible tools stay denied. A prettier demo schema is not a week.
When is this not worth doing yet?
If the job is a fixed lookup-then-write flowchart, ship a workflow with one strict schema on the write node instead of an agent catalog. If you have no side-effecting tools — no send, no hold, no refund — you do not have a schema-quality problem yet; you have a chatbot. If nobody will run the overlap test or own the enum when the API adds a value, do not allowlist production writes.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot