Spurlock Studios
Contact
Share LinkedIn X
A scuffed work smartphone with a blank glowing circular button. Thesis: AGENT KEEP CALLING SAME FAILED.

Your agent keeps calling the same failed tool because the runner treats every model turn as progress. The tool already returned an error. The model still proposes the same call. Nothing fingerprints that retry, nothing caps same-tool executes, and auth or policy failures never escalate — they grind. A longer prompt will not invent those brakes.

This spoke sits under the Agentic Systems Operating Manual. It is the operator diagnosis: why your run will not stop. Detector internals live on why agents loop on failed tools. If the first live week looked like a demo that died on real APIs, read why agent demos fail in production.

The short answer

  • Same tool plus same normalized args is one attempt. Repeating it is not diligence.
  • Fingerprint the call before execute. Cap how many times that fingerprint may run. Cap the tool name as a backstop.
  • Auth (401 / 403) and policy refuses escalate on the first hit. The model cannot fix credentials or a tenant gate.
  • Transient classes (429, timeout, some 5xx) may retry with a clock and a small cap. Permanent classes stop.
  • I will not mint a studio-wide “loop rate.” Yours is the share of runs that hit the same fingerprint twice. Measure that.
GuardrailWhat it stopsWhat it misses if used alone
Fingerprint of tool + argsIdentical 404 hammerNearby invented ids on the same tool
Same-tool call capThrash across similar argsA wander onto a second write tool
Auth / policy escalateCredential and tenant grindA missing-record loop that is not auth
Turn budgetEndless stepsSame fingerprint on step 3 of 8

Ship the first three this week. The turn budget is extra, not a substitute.

What does the same failed tool look like from the ops seat?

You are not reading a paper on no-progress detection. You are watching a run. The same tool name keeps showing up. The args look familiar. The error text is already in the log. The model writes another “let me try that again.”

That is the symptom. Name it before you blame the model.

What you seeWhat it usually isWhat it is not
Same tool, same id, same error, no waitIdentical fingerprint grind“Being thorough”
Same tool, nearby ids after a 404Arg wander after a permanent missA new strategy
Same tool, 401 / 403, immediate retryAuth grindA flaky network
Same tool, 429, immediate retryMissing clockPersistence
Two tools bouncing with no new artifactHandoff or alternate-tool thrashA same-tool loop

Open the last failed run and walk this list before you edit a prompt:

  1. Write down the tool name on every execute.
  2. Write down the args a human would call “the same call” (id, filter, payload keys). Strip request ids and timestamps.
  3. Write down the error class: auth, policy, not-found, validation, rate-limit, timeout, 5xx, empty result, unknown.
  4. Count how many executes share tool + normalized args.
  5. Note whether anything waited (Retry-After, backoff) between those executes.

Four of five of these is enough to call it a same-failed-tool loop:

  • Same tool name
  • Same normalized args
  • Same error class, or empty result treated as “try again”
  • No backoff
  • No new state you could show a reviewer

If a human watching the log would say “it is doing the same thing again,” the harness should already have refused the execute. You should not need a token budget to notice.

Why does the model keep proposing a call that already failed?

Models optimize for finishing the user job. An error string is more context, not a stop. Unless the harness refuses the next execute, the next plan is often “call it again.” That is rational under incomplete control. It is expensive under write tools and paid APIs.

After 20,000+ hours on agentic systems and 500+ automations, the boring version is this: the model is not uniquely stubborn. The cage never told the runner the attempt was terminal.

Fuel you will actually seeWhat the model concludesWhat you should enforce
"failed" with no classMaybe a blipClassify; unknown fails closed
Empty [] on the same queryMaybe a bad search — try the same searchCount empty-same-query as a hit
“Keep going until done” in the promptPermission to grindIgnore; caps still fire
History summarized, failure droppedA fresh planInject a failure ledger every turn
Tool error returned only to the model“I can recover”Harness refuse on the next identical propose

Anthropic’s context engineering note treats agents as models using tools in a loop, with a limited attention budget. If the only copy of “this id already 404’d” lives in chat tokens, a compress will drop it and the loop returns. That is an operator problem: your evidence lived in the wrong store.

Do not moralize the model. Instrument the refuse path.

  • Tool errors have a machine-readable class, not only a sentence
  • The runner can deny an execute the model asked for
  • Failure evidence lives outside the prompt
  • “Helpful” prompt edits cannot disable the deny

If any box is empty, you still have a suggestion engine with an API key.

How do I fingerprint retries before they execute?

A fingerprint is how you stop arguing about “same.” Compute it before the HTTP client, not after the bill.

Operator version, before you write a hash function: if you would paste two arg blobs into a diff and call them the same attempt, they are the same fingerprint.

Include in the fingerprintExclude
Tool nameBearer tokens, cookies, Authorization
Canonical args (sorted keys)Injected request_id, harness timestamps
Side-effect class: read / writeVolatile pagination cursors you did not send
Tenant / resource id the call is aboutTrace ids the runner attached

Normalization rules you can enforce this week without a new framework:

  1. Sort object keys.
  2. Trim string ids. Lowercase only where the downstream system is actually case-insensitive.
  3. Treat omitted and null the same way your tool schema treats them — pick one rule per tool and write it down.
  4. Strip fields the harness injected. Hashing your own noise creates fake “new” calls.
  5. Empty result on the same fingerprint still counts. [] is a result.

Store per run, even if the store is a JSON file beside the trace:

FieldWhy the operator needs it
fingerprintIdentity of the attempt
toolFor the same-tool cap
countHow many times this run hit it
error_classAuth vs not-found vs transient
blockedDid the runner refuse
last_atDid anything wait

Policy you can ship as numbers, not vibes (tune per job type; these are starting caps, not a benchmark):

SituationAction
Same fingerprint + auth / policy / not-found / validationRefuse immediately; escalate
Same fingerprint + rate-limit / timeoutAllow up to 2 waits with a clock, then refuse
Same fingerprint + unknownRefuse in production
Same tool, new args, no state changeSoft warn; hard cap the tool name

The model may propose the call. The harness decides whether it runs. Fingerprinting is that decision’s input.

  • Pre-execute fingerprint exists for every allowlisted tool
  • Normalize has a unit test per tool, including “same id, extra whitespace”
  • Block path never reaches the vendor client
  • Writes after blocked are a bug, not a log line

How do I cap same-tool calls per run?

Fingerprints stop the identical hammer. Operators still get burned by the next species: the agent 404s on c_1842, then tries c_1843 through c_1850. Each fingerprint is new. The tool is not.

That is why you cap same-tool calls, not only identical args.

CapWhat it isUse it for
Same-fingerprint capMax executes of one normalized call404 / validation hammer
Same-tool capMax executes of that tool name in the runNearby-id thrash, empty-search spam
Turn capMax model stepsRunaway plans that switch tools
Spend / row capMax dollars or rows touchedPaid enrichment, bulk writes

Illustrative starting numbers — measure your successful trajectories and replace these. Do not treat them as a published loop rate.

Job typeSame-fingerprintSame-toolNotes
Read-only lookup1 if permanent; 3 if retryable4Empty result counts
Single CRM write1 on permanent2Second write after a block is an incident
Multi-step research2 per fingerprint8Still not “until the budget dies”
Human-in-the-loop11 after a refuseEscalate, do not grind

Procedure to install the tool cap without waiting on a rewrite:

  1. List tools in the allowlist. Mark each read or write.
  2. Put a counter on the run record keyed by tool name.
  3. Increment on execute, not on propose. A blocked propose should not consume the cap — or you will starve a legitimate alternate.
  4. When the cap hits, refuse with a reason the ops screen can group (tool_cap, or your catalog equivalent).
  5. On write tools, the cap is lower. Reads looping are a bill. Writes looping are a mess.
TemptationWhy it fails
One global “max tool calls = 50”A research job and a CRM patch are not the same job
Cap proposes, not executesThe model burns the budget on blocked ideas
Cap only writesEmpty-search loops still cost tokens and vendor reads
Raise the cap when a job is “hard”Hard jobs need escalate, not a longer hammer

A same-tool cap is a backstop. It is not permission to skip fingerprints.

When should I escalate on auth or policy errors?

Auth and policy are not model problems. Retrying them is how you turn a missing secret into a lockout, and a tenant gate into a denial-of-wallet.

RFC 9110 §15.5.2 defines 401 Unauthorized as: the request was not applied because it lacks valid authentication credentials. The user agent may repeat with a new Authorization header. Your agent does not have a new header. It has the same runner identity. Repeating the same credentials is not a new attempt.

RFC 9110 §15.5.4 is blunter on 403 Forbidden: the server understood the request and refuses to fulfill it. If credentials were provided, they were insufficient. The client should not automatically repeat with the same credentials. A request might be forbidden for reasons unrelated to credentials. Either way, the model cannot talk the server into a scope it does not have.

SignalFirst actionWho you pageRetry?
401 / missing or rejected tokenTerminate; attach redacted auth classWhoever owns secrets / rotationNo
403 / insufficient scopeTerminate; attach required vs granted scope (no secret)Whoever owns rolesNo
Policy gate refuse (wrong tenant, write freeze, spend cap)Terminate policy_blockedOps / policy ownerNo
404 / not foundTerminate or one alternate read toolWhoever owns the idNo on the same id
422 / schema rejectTerminate; attach schema pathWhoever owns the tool contractNo
429 + Retry-AfterWait, cappedNobody unless the cap tripsYes, with a clock
Timeout / some 5xxBackoff, cappedVendor if it persistsYes, then escalate

MDN’s HTTP authentication guide makes the same split operators forget: 401 is “I do not know who you are”; 403 is “I know who you are and you still cannot do this.” Browsers do not keep prompting after a real 403. Your agent should not either.

OWASP LLM06:2025 Excessive Agency names the three fuels: too much functionality, too much permission, too much autonomy. A retry loop on 403 is excessive autonomy with the original over-permission still attached. The mitigation is not a sterner system prompt. It is complete mediation in the downstream system, plus a human path for high-impact actions.

Escalate package (minimum):

  1. Terminal reason (tool_auth_error, policy_blocked, or your catalog).
  2. Tool name, fingerprint, error class.
  3. Redacted scope — never the secret.
  4. Last artifact the run produced.
  5. The human action: rotate a token, grant a role, fix an id, or kill the job.
  • Auth classes map to terminate, not to “try the tool again”
  • Policy refuses skip the model’s next plan
  • Secrets never land in the escalate ticket
  • Someone is named as the owner of creds vs roles vs records

If you cannot name the owner, you do not have an escalate path. You have a chat window.

How do I tell a legitimate retry from a grind?

Ops needs a crisp split. Without it, every retry looks like care and every stop looks like you gave up.

SignalLegitimate retryGrind
ToolSame or a real alternateSame
ArgsChanged in a way that could succeed (new page, new id you have evidence for)Identical, or guessed neighbors after a permanent miss
ClockBackoff / Retry-AfterImmediate hammer
Error classTransientPermanent or policy
HarnessNew fingerprint, or marked retryable under a capSame fingerprint, or same tool past the cap

Illustrative — not a client result. Job: patch contact c_1842. Tool returns 404, not retryable. Without a fingerprint and a tool cap, the run hits the same id, then nearby ids, then a write-shaped “create if missing” if that tool exists. With both: first failure records the fingerprint; second identical propose is refused; same-tool cap stops the neighbor walk; the run escalates with the real id problem.

A legitimate sibling on the same tool: 429, Retry-After: 8, same fingerprint, wait, one more execute, success. That is diligence because the flag and the clock said so.

PatternDetectionOperator fix
Same-tool identical argsFingerprint countBlock; escalate if auth/policy/not-found
Same-tool new args, no progressSame-tool cap + state checksumBlock the tool; do not invent ids
Alternate-tool thrashTool-set cycle without artifact changeRequire a checksum to switch tools
ABAB handoffPackage hash / cycleBreak the cycle; do not stretch a tool cap to cover a graph

State checksum, in operator language: a hash of the artifact you would show a reviewer (record body, draft, ticket fields). If the checksum did not move, the last tool call was not progress, even if the name changed.

  • You can point at a row and say legitimate vs grind in one sentence
  • Transient retries have a clock
  • Permanent retries have a refuse
  • Neighbor-id walks have a tool cap

If you cannot draw that table on a whiteboard, you will keep arguing in Slack while the run spends.

What breaks when a write tool is in that loop?

Reads looping are a bill and a noisy vendor. Writes looping are records you then have to unwind.

Failure mode, named: write hammer after a permanent miss. The lookup 404s. The agent still has create_contact or an upsert. It “helps.” You now have junk rows, duplicate customers, or a patch applied to the wrong tenant if the id was guessed.

OWASP LLM10:2025 Unbounded Consumption is usually framed as attackers burning inference. The operator cousin is an unbounded tool loop: each extra execute is another model turn plus another vendor call. I will not invent a percent of runs this happens on. The cost is mechanical: unbounded executes times unit price, plus cleanup.

Write-loop speciesWhat it costsBrake
Identical write after 404Junk create, or twelve no-ops you still pay forFingerprint + retryable: false
Neighbor-id writesWrong-record damageSame-tool cap
Auth retry on a writeLockouts, SIEM noise, possible account flagsEscalate on first 401/403
Policy retry on a freezeProof the freeze is advisoryFreeze in the runner, not the prompt
Empty-search then blast-createInvented entitiesEmpty result counts; create is gated

What you do instead:

  1. Mark write tools in the catalog. They get the tightest fingerprint and tool caps.
  2. Forbid create-on-404 unless a named state allows it and a human has already approved the id class.
  3. Dry-run the write path in week one: “would write X.” Dual-write to a shadow field before hard writes.
  4. After any block, assert zero write executes in the rest of the run.
  5. Keep an unwind note in the escalate package: which row, which field, which timestamp.
  • Write tools have a side-effect class in the catalog
  • Create is not the default recovery from not-found
  • Post-block write count is a monitored zero
  • Someone can undo the last agent write from the escalate ticket

A prompt that says “be careful with writes” is not a brake. The runner refusing create_contact after get_contact 404’d is a brake.

How do I read a trace when the agent will not stop?

If you cannot see the repeated fingerprint, you will debug the model. Debug the execute list first.

QuestionWhere to lookIf you cannot answer
Which tool fired?Execute spans / log lines, not chat proseYou do not have traces
Same args or new args?Normalized arg hash, or a redacted args panelYou cannot fingerprint yet
Error class?Status code / error_code, not the model’s summaryYou threw away HTTP
Did the harness refuse?A blocked flagYou only log after success
Did anything wait?Timestamps between executesYou cannot tell grind from backoff
Why did the run end?Terminal reason codeYou will not trend this next week

Operator pass on a stuck run (ten minutes, one run):

  1. Export executes in order: time, tool, arg hash, status, retryable, blocked.
  2. Group by arg hash. Any group with count ≥ 2 on a permanent class is the incident.
  3. Group by tool name. Any tool past its cap is the backstop catching wander.
  4. Check the last reason code. If it is generic error or missing, fix the catalog before you fix the prompt.
  5. Check whether a summarize step sits between attempt 1 and attempt 2. If yes, confirm the ledger survived.
Metric to start this monthWhyHedge
Share of runs with any fingerprint count ≥ 2Loop pressureYour baseline, not a published industry rate
Share terminated tool_auth_error / policy_blockedAre you escalating or grindingNeeds a reason catalog
Executes after blockedBrake bypassShould be zero
Cost of runs with fingerprint ≥ 2 vs the restMoney on the floorUnit cost on your tools

Do not wait for a weekly CSV to learn you hammered a 404 overnight. One alert on “blocked-fingerprint rate after a deploy” is enough to start. The complementary detector post is where the catalog of reason codes is specified; this page is where you use the trace as an operator.

When is a workflow enough instead of an agent?

If the honest path is “look up this id; if it exists, patch these fields; if it 404s, stop,” you do not need a model choosing tools in a loop. You need a workflow with an error branch.

Anthropic’s building-effective-agents essay draws the split: workflows follow predefined code paths; agents let the model direct process and tool use. They tell teams to find the simplest solution and add complexity only when it demonstrably helps. A same-failed-tool loop is often proof you added a director to a job that needed a branch.

Job shapePreferWhy
Known id → known write → stop on 404WorkflowThe miss is a branch, not a plan
Auth failure → page secretsWorkflow / runnerNo model step adds a valid token
Policy refuse → stopRunner gateThe model should not see a workaround
Ambiguous record match across systemsAgent, with capsYou need a search strategy — still capped
Open research with a write at the endHybrid: agent reads, workflow writesDo not let research thrash a write tool

Decision list:

  1. Can you draw the happy path and the 404 path without a model? Ship a workflow.
  2. Does the job need a search strategy under uncertainty? An agent may earn the loop — with fingerprint, tool cap, and escalate.
  3. Is the failure class auth or policy? Neither topology should retry. The runner stops.
  4. Are you using an agent because the demo looked smarter? Read the production-control spoke before you widen autonomy.

Guardrails you need either way:

  • Typed tool schemas
  • Tenant bound at the runner
  • Fingerprint + same-tool cap on every tool the job can call
  • Auth/policy escalate with an owner
  • A terminal other than “the model said done”

An agent is not a retry engine. If retries are the product, you wanted a queue with backoff.

What should I not do while it is still looping?

Add “please don’t retry failed tools” to the system prompt as the only fix. It will drift on the next “be more persistent” edit.

Retry every error three times. Auth, policy, validation, and not-found are not transient. HTTP already sorted this. Your catalog should too.

Raise temperature, swap models, or buy a larger context window. Irrelevant to a deterministic 404 with the same args.

Log the loop and keep executing. Observation without a refuse path is a museum exhibit.

Raise the same-tool cap because the job is “hard.” Hard jobs escalate with evidence.

Treat an empty search as progress. [] on the same query is a hit.

Let the model pick a second write tool after the first 404. That is a workaround, not recovery.

TemptationWhy it failsDo this instead
Longer promptDrift + summarizationFingerprint gate
Retry-everythingPermanent classes hammerClassify; fail closed on unknown
New modelSame missing brakeSame harness
Unlimited turnsBill with no artifactFinite cap + escalate
Hide 403 as 404 and keep searchingYou deleted the escalate signalKeep the status; page roles
“Just this once” disable the cap in prodThe exception becomes the pathChange the job type’s numbers in config, in staging first

OWASP’s unbounded-consumption entry is the cost framing: uncontrolled inference (and, for you, uncontrolled tool executes) is how a pay-per-use system becomes a surprise invoice. Caps are not pessimism. They are how the job stays a job.

How do I measure this in production?

You asked how to evaluate it. Do not wait for a vendor “agent reliability” score. Count repeats.

QuestionMetricPass shape (set yours)
Is the identical hammer dying?Fingerprint count ≥ 2 on permanent classesFalls after the gate ships; never a claimed industry %
Is neighbor-id thrash dying?Same-tool executes per runMedian at or under the cap
Are we escalating auth?Runs with 401/403 that still executed againZero
Are we escalating policy?Policy refuses that still hit the vendorZero
Did the brake hold after a prompt edit?Golden case for a known bad idStill red on second execute
Are writes safe?Write executes after blockedZero

Procedure for a baseline this week — no invented loop rate, just your numbers:

  1. Take the last 20 failed runs (or all of last week if you have fewer).
  2. For each, mark: identical-fingerprint repeat, same-tool wander, auth retry, policy retry, legitimate backoff, other.
  3. Write the counts on one page. That page is the baseline.
  4. Ship fingerprint + tool cap + auth/policy terminate.
  5. Repeat the 20-run sample after the change. Compare counts, not adjectives.
You will be tempted to reportReport this instead
“Loops are down a lot”Identical-fingerprint repeats: 11/20 → 2/20 on this sample
“The new model is better”Same sample, same harness, one variable
“Zero loops”Zero identical repeats; wander still has a tool-cap metric
A percent copied from a vendor blogYour sample, dated, with the definition you used

Golden-case the known offender so the measurement is not only a retro:

  • Stub the permanent error
  • Assert the second identical execute never hits the client
  • Assert auth/policy never hits the client twice
  • Assert the same-tool cap fires on a neighbor walk
  • Force a summarize between attempts; assert the block still holds

If a prompt edit that “makes the agent more persistent” does not turn that case red, you are still measuring vibes.

What can I ship in a week?

You asked what to skip if you only have a week. Skip the dashboard redesign, the multi-agent oscillation work, and the prompt rewrite. Keep the refuse path.

Day-by-day, one job, write tools first:

  1. Day 1. Dump executes from five real failures. Classify error classes by hand. Name the identical hammer vs the wander.
  2. Day 2. Normalize + fingerprint for those tools. Unit test “same id, different key order.”
  3. Day 3. Pre-execute refuse on fingerprint. Auth/policy terminate on first hit. Unknown fails closed.
  4. Day 4. Same-tool cap on the write tool. Reason code on the run record. Escalate package template.
  5. Day 5. Golden case + a 20-run recount. If writes still execute after block, that is the only bug you work.
If you only have two daysDoSkip
Do firstFingerprint refuse + auth/policy terminatePretty traces
NextSame-tool cap on writesAlternate-tool graphs
ProofOne golden bad-id caseFleet-wide percentiles

Checklist for the PR that actually ships:

  • Fingerprint computed pre-execute
  • Same-tool counter on the run
  • 401 / 403 / policy refuse never retry
  • Reason code is an enum
  • Escalate names an owner
  • Golden case green
  • Prompt text is not the brake

If you only have a day: fingerprint the worst write tool, cap it at 2, terminate auth. Coarse beats a better paragraph.

What does an agentic pilot install for this?

In a Spurlock Studios $1,500 · 5-day agentic pilot, same-failed-tool brakes are part of the thin harness for one job: fingerprints on write-capable tools, a same-tool cap, auth/policy terminate, a reason code, and at least one golden bad-id case. Full fleet dashboards and multi-agent cycle detection can wait. The identical hammer should not.

You leave with counts on your last sample, not a promised loop-rate. Architecture context stays in the operating manual. Start from /agentic.

Pilot includesPilot does not include
One job, real tools, real error classesA claim that loops cannot happen
Fingerprint + same-tool cap + escalateA new model as the fix
Golden case in CIPrompt-only “don’t retry” language as the control
Reason codes you can trend laterA full observability rebuild

If the job is still a known branch, the pilot may tell you to ship a workflow. That is a successful week.

FAQ

Why does my agent keep calling the same failed tool?

Because the runner treats every model turn as progress and never fingerprints the retry or caps same-tool executes. The model sees an error and proposes the same call; nothing refuses the next execute. Auth and policy failures need an escalate path on the first hit, not another attempt with the same credentials.

How do I measure whether a same-tool retry cap is working?

Count identical-fingerprint repeats and same-tool executes on a dated sample of failed runs, then recount after the gate ships. Auth and policy classes should show zero second executes. Do not publish a vendor-style loop rate — report your sample, your definition, and whether the golden bad-id case still blocks.

What usually fails first when teams try this?

Error classification. Teams retry 401, 403, and policy refuses as if they were timeouts, or they fingerprint after the HTTP client so the bill already happened. The second miss is skipping the same-tool cap, so a 404 on one id becomes a neighbor walk that still looks like “progress” in the trace.

How long does this take to show results?

A coarse fingerprint plus a write-tool cap plus auth terminate can show up on the next failing run — often inside a week if you already have execute logs. Trendable reason codes take longer because you need a catalog and a sample. Do not wait on a dashboard to refuse the second identical call.

What should I skip if I only have a week?

Skip prompt rewrites, model swaps, multi-agent oscillation work, and a full observability rebuild. Fingerprint the worst write tool, cap that tool, terminate auth and policy on first hit, and add one golden bad-id case. That is the refuse path. Everything else is furniture.

When is this not worth doing yet?

If the job is a known lookup-then-stop path, do not add an agent so you can then cap it — ship a workflow with an error branch. If you cannot list the tools or see executes, instrument one run before you write a fingerprint function. If writes are already unconstrained in production, freeze writes first; then add retry brakes.

CTA

Cap the retry. Escalate the auth error. Then decide if the job still needs an agent.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

Why does my agent keep calling the same failed tool?
Because the runner treats every model turn as progress and never fingerprints the retry or caps same-tool executes. The model sees an error and proposes the same call; nothing refuses the next execute. Auth and policy failures need an escalate path on the first hit, not another attempt with the same credentials.
How do I measure whether a same-tool retry cap is working?
Count identical-fingerprint repeats and same-tool executes on a dated sample of failed runs, then recount after the gate ships. Auth and policy classes should show zero second executes. Do not publish a vendor-style loop rate — report your sample, your definition, and whether the golden bad-id case still blocks.
What usually fails first when teams try this?
Error classification. Teams retry `401`, `403`, and policy refuses as if they were timeouts, or they fingerprint after the HTTP client so the bill already happened. The second miss is skipping the same-tool cap, so a 404 on one id becomes a neighbor walk that still looks like “progress” in the trace.
How long does this take to show results?
A coarse fingerprint plus a write-tool cap plus auth terminate can show up on the next failing run — often inside a week if you already have execute logs. Trendable reason codes take longer because you need a catalog and a sample. Do not wait on a dashboard to refuse the second identical call.
What should I skip if I only have a week?
Skip prompt rewrites, model swaps, multi-agent oscillation work, and a full observability rebuild. Fingerprint the worst write tool, cap that tool, terminate auth and policy on first hit, and add one golden bad-id case. That is the refuse path. Everything else is furniture.
When is this not worth doing yet?
If the job is a known lookup-then-stop path, do not add an agent so you can then cap it — ship a workflow with an error branch. If you cannot list the tools or see executes, instrument one run before you write a fingerprint function. If writes are already unconstrained in production, freeze writes first; then add retry brakes.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot