Why does my agent keep calling the same failed tool
Your agent retries one failed tool because nothing fingerprints the call or caps same-tool executes. Escalate auth and policy on first hit — do not grind.
William Spurlock Founder — Spurlock Studios 25 MIN
Your agent keeps calling the same failed tool because the runner treats every model turn as progress. The tool already returned an error. The model still proposes the same call. Nothing fingerprints that retry, nothing caps same-tool executes, and auth or policy failures never escalate — they grind. A longer prompt will not invent those brakes.
This spoke sits under the Agentic Systems Operating Manual. It is the operator diagnosis: why your run will not stop. Detector internals live on why agents loop on failed tools. If the first live week looked like a demo that died on real APIs, read why agent demos fail in production.
The short answer
- Same tool plus same normalized args is one attempt. Repeating it is not diligence.
- Fingerprint the call before execute. Cap how many times that fingerprint may run. Cap the tool name as a backstop.
- Auth (
401/403) and policy refuses escalate on the first hit. The model cannot fix credentials or a tenant gate. - Transient classes (
429, timeout, some 5xx) may retry with a clock and a small cap. Permanent classes stop. - I will not mint a studio-wide “loop rate.” Yours is the share of runs that hit the same fingerprint twice. Measure that.
| Guardrail | What it stops | What it misses if used alone |
|---|---|---|
| Fingerprint of tool + args | Identical 404 hammer | Nearby invented ids on the same tool |
| Same-tool call cap | Thrash across similar args | A wander onto a second write tool |
| Auth / policy escalate | Credential and tenant grind | A missing-record loop that is not auth |
| Turn budget | Endless steps | Same fingerprint on step 3 of 8 |
Ship the first three this week. The turn budget is extra, not a substitute.
What does the same failed tool look like from the ops seat?
You are not reading a paper on no-progress detection. You are watching a run. The same tool name keeps showing up. The args look familiar. The error text is already in the log. The model writes another “let me try that again.”
That is the symptom. Name it before you blame the model.
| What you see | What it usually is | What it is not |
|---|---|---|
| Same tool, same id, same error, no wait | Identical fingerprint grind | “Being thorough” |
| Same tool, nearby ids after a 404 | Arg wander after a permanent miss | A new strategy |
Same tool, 401 / 403, immediate retry | Auth grind | A flaky network |
Same tool, 429, immediate retry | Missing clock | Persistence |
| Two tools bouncing with no new artifact | Handoff or alternate-tool thrash | A same-tool loop |
Open the last failed run and walk this list before you edit a prompt:
- Write down the tool name on every execute.
- Write down the args a human would call “the same call” (id, filter, payload keys). Strip request ids and timestamps.
- Write down the error class: auth, policy, not-found, validation, rate-limit, timeout, 5xx, empty result, unknown.
- Count how many executes share tool + normalized args.
- Note whether anything waited (
Retry-After, backoff) between those executes.
Four of five of these is enough to call it a same-failed-tool loop:
- Same tool name
- Same normalized args
- Same error class, or empty result treated as “try again”
- No backoff
- No new state you could show a reviewer
If a human watching the log would say “it is doing the same thing again,” the harness should already have refused the execute. You should not need a token budget to notice.
Why does the model keep proposing a call that already failed?
Models optimize for finishing the user job. An error string is more context, not a stop. Unless the harness refuses the next execute, the next plan is often “call it again.” That is rational under incomplete control. It is expensive under write tools and paid APIs.
After 20,000+ hours on agentic systems and 500+ automations, the boring version is this: the model is not uniquely stubborn. The cage never told the runner the attempt was terminal.
| Fuel you will actually see | What the model concludes | What you should enforce |
|---|---|---|
"failed" with no class | Maybe a blip | Classify; unknown fails closed |
Empty [] on the same query | Maybe a bad search — try the same search | Count empty-same-query as a hit |
| “Keep going until done” in the prompt | Permission to grind | Ignore; caps still fire |
| History summarized, failure dropped | A fresh plan | Inject a failure ledger every turn |
| Tool error returned only to the model | “I can recover” | Harness refuse on the next identical propose |
Anthropic’s context engineering note treats agents as models using tools in a loop, with a limited attention budget. If the only copy of “this id already 404’d” lives in chat tokens, a compress will drop it and the loop returns. That is an operator problem: your evidence lived in the wrong store.
Do not moralize the model. Instrument the refuse path.
- Tool errors have a machine-readable class, not only a sentence
- The runner can deny an execute the model asked for
- Failure evidence lives outside the prompt
- “Helpful” prompt edits cannot disable the deny
If any box is empty, you still have a suggestion engine with an API key.
How do I fingerprint retries before they execute?
A fingerprint is how you stop arguing about “same.” Compute it before the HTTP client, not after the bill.
Operator version, before you write a hash function: if you would paste two arg blobs into a diff and call them the same attempt, they are the same fingerprint.
| Include in the fingerprint | Exclude |
|---|---|
| Tool name | Bearer tokens, cookies, Authorization |
| Canonical args (sorted keys) | Injected request_id, harness timestamps |
| Side-effect class: read / write | Volatile pagination cursors you did not send |
| Tenant / resource id the call is about | Trace ids the runner attached |
Normalization rules you can enforce this week without a new framework:
- Sort object keys.
- Trim string ids. Lowercase only where the downstream system is actually case-insensitive.
- Treat omitted and
nullthe same way your tool schema treats them — pick one rule per tool and write it down. - Strip fields the harness injected. Hashing your own noise creates fake “new” calls.
- Empty result on the same fingerprint still counts.
[]is a result.
Store per run, even if the store is a JSON file beside the trace:
| Field | Why the operator needs it |
|---|---|
fingerprint | Identity of the attempt |
tool | For the same-tool cap |
count | How many times this run hit it |
error_class | Auth vs not-found vs transient |
blocked | Did the runner refuse |
last_at | Did anything wait |
Policy you can ship as numbers, not vibes (tune per job type; these are starting caps, not a benchmark):
| Situation | Action |
|---|---|
| Same fingerprint + auth / policy / not-found / validation | Refuse immediately; escalate |
| Same fingerprint + rate-limit / timeout | Allow up to 2 waits with a clock, then refuse |
| Same fingerprint + unknown | Refuse in production |
| Same tool, new args, no state change | Soft warn; hard cap the tool name |
The model may propose the call. The harness decides whether it runs. Fingerprinting is that decision’s input.
- Pre-execute fingerprint exists for every allowlisted tool
- Normalize has a unit test per tool, including “same id, extra whitespace”
- Block path never reaches the vendor client
- Writes after
blockedare a bug, not a log line
How do I cap same-tool calls per run?
Fingerprints stop the identical hammer. Operators still get burned by the next species: the agent 404s on c_1842, then tries c_1843 through c_1850. Each fingerprint is new. The tool is not.
That is why you cap same-tool calls, not only identical args.
| Cap | What it is | Use it for |
|---|---|---|
| Same-fingerprint cap | Max executes of one normalized call | 404 / validation hammer |
| Same-tool cap | Max executes of that tool name in the run | Nearby-id thrash, empty-search spam |
| Turn cap | Max model steps | Runaway plans that switch tools |
| Spend / row cap | Max dollars or rows touched | Paid enrichment, bulk writes |
Illustrative starting numbers — measure your successful trajectories and replace these. Do not treat them as a published loop rate.
| Job type | Same-fingerprint | Same-tool | Notes |
|---|---|---|---|
| Read-only lookup | 1 if permanent; 3 if retryable | 4 | Empty result counts |
| Single CRM write | 1 on permanent | 2 | Second write after a block is an incident |
| Multi-step research | 2 per fingerprint | 8 | Still not “until the budget dies” |
| Human-in-the-loop | 1 | 1 after a refuse | Escalate, do not grind |
Procedure to install the tool cap without waiting on a rewrite:
- List tools in the allowlist. Mark each
readorwrite. - Put a counter on the run record keyed by tool name.
- Increment on execute, not on propose. A blocked propose should not consume the cap — or you will starve a legitimate alternate.
- When the cap hits, refuse with a reason the ops screen can group (
tool_cap, or your catalog equivalent). - On write tools, the cap is lower. Reads looping are a bill. Writes looping are a mess.
| Temptation | Why it fails |
|---|---|
| One global “max tool calls = 50” | A research job and a CRM patch are not the same job |
| Cap proposes, not executes | The model burns the budget on blocked ideas |
| Cap only writes | Empty-search loops still cost tokens and vendor reads |
| Raise the cap when a job is “hard” | Hard jobs need escalate, not a longer hammer |
A same-tool cap is a backstop. It is not permission to skip fingerprints.
When should I escalate on auth or policy errors?
Auth and policy are not model problems. Retrying them is how you turn a missing secret into a lockout, and a tenant gate into a denial-of-wallet.
RFC 9110 §15.5.2 defines 401 Unauthorized as: the request was not applied because it lacks valid authentication credentials. The user agent may repeat with a new Authorization header. Your agent does not have a new header. It has the same runner identity. Repeating the same credentials is not a new attempt.
RFC 9110 §15.5.4 is blunter on 403 Forbidden: the server understood the request and refuses to fulfill it. If credentials were provided, they were insufficient. The client should not automatically repeat with the same credentials. A request might be forbidden for reasons unrelated to credentials. Either way, the model cannot talk the server into a scope it does not have.
| Signal | First action | Who you page | Retry? |
|---|---|---|---|
401 / missing or rejected token | Terminate; attach redacted auth class | Whoever owns secrets / rotation | No |
403 / insufficient scope | Terminate; attach required vs granted scope (no secret) | Whoever owns roles | No |
| Policy gate refuse (wrong tenant, write freeze, spend cap) | Terminate policy_blocked | Ops / policy owner | No |
404 / not found | Terminate or one alternate read tool | Whoever owns the id | No on the same id |
422 / schema reject | Terminate; attach schema path | Whoever owns the tool contract | No |
429 + Retry-After | Wait, capped | Nobody unless the cap trips | Yes, with a clock |
| Timeout / some 5xx | Backoff, capped | Vendor if it persists | Yes, then escalate |
MDN’s HTTP authentication guide makes the same split operators forget: 401 is “I do not know who you are”; 403 is “I know who you are and you still cannot do this.” Browsers do not keep prompting after a real 403. Your agent should not either.
OWASP LLM06:2025 Excessive Agency names the three fuels: too much functionality, too much permission, too much autonomy. A retry loop on 403 is excessive autonomy with the original over-permission still attached. The mitigation is not a sterner system prompt. It is complete mediation in the downstream system, plus a human path for high-impact actions.
Escalate package (minimum):
- Terminal reason (
tool_auth_error,policy_blocked, or your catalog). - Tool name, fingerprint, error class.
- Redacted scope — never the secret.
- Last artifact the run produced.
- The human action: rotate a token, grant a role, fix an id, or kill the job.
- Auth classes map to terminate, not to “try the tool again”
- Policy refuses skip the model’s next plan
- Secrets never land in the escalate ticket
- Someone is named as the owner of creds vs roles vs records
If you cannot name the owner, you do not have an escalate path. You have a chat window.
How do I tell a legitimate retry from a grind?
Ops needs a crisp split. Without it, every retry looks like care and every stop looks like you gave up.
| Signal | Legitimate retry | Grind |
|---|---|---|
| Tool | Same or a real alternate | Same |
| Args | Changed in a way that could succeed (new page, new id you have evidence for) | Identical, or guessed neighbors after a permanent miss |
| Clock | Backoff / Retry-After | Immediate hammer |
| Error class | Transient | Permanent or policy |
| Harness | New fingerprint, or marked retryable under a cap | Same fingerprint, or same tool past the cap |
Illustrative — not a client result. Job: patch contact c_1842. Tool returns 404, not retryable. Without a fingerprint and a tool cap, the run hits the same id, then nearby ids, then a write-shaped “create if missing” if that tool exists. With both: first failure records the fingerprint; second identical propose is refused; same-tool cap stops the neighbor walk; the run escalates with the real id problem.
A legitimate sibling on the same tool: 429, Retry-After: 8, same fingerprint, wait, one more execute, success. That is diligence because the flag and the clock said so.
| Pattern | Detection | Operator fix |
|---|---|---|
| Same-tool identical args | Fingerprint count | Block; escalate if auth/policy/not-found |
| Same-tool new args, no progress | Same-tool cap + state checksum | Block the tool; do not invent ids |
| Alternate-tool thrash | Tool-set cycle without artifact change | Require a checksum to switch tools |
| ABAB handoff | Package hash / cycle | Break the cycle; do not stretch a tool cap to cover a graph |
State checksum, in operator language: a hash of the artifact you would show a reviewer (record body, draft, ticket fields). If the checksum did not move, the last tool call was not progress, even if the name changed.
- You can point at a row and say legitimate vs grind in one sentence
- Transient retries have a clock
- Permanent retries have a refuse
- Neighbor-id walks have a tool cap
If you cannot draw that table on a whiteboard, you will keep arguing in Slack while the run spends.
What breaks when a write tool is in that loop?
Reads looping are a bill and a noisy vendor. Writes looping are records you then have to unwind.
Failure mode, named: write hammer after a permanent miss. The lookup 404s. The agent still has create_contact or an upsert. It “helps.” You now have junk rows, duplicate customers, or a patch applied to the wrong tenant if the id was guessed.
OWASP LLM10:2025 Unbounded Consumption is usually framed as attackers burning inference. The operator cousin is an unbounded tool loop: each extra execute is another model turn plus another vendor call. I will not invent a percent of runs this happens on. The cost is mechanical: unbounded executes times unit price, plus cleanup.
| Write-loop species | What it costs | Brake |
|---|---|---|
Identical write after 404 | Junk create, or twelve no-ops you still pay for | Fingerprint + retryable: false |
| Neighbor-id writes | Wrong-record damage | Same-tool cap |
| Auth retry on a write | Lockouts, SIEM noise, possible account flags | Escalate on first 401/403 |
| Policy retry on a freeze | Proof the freeze is advisory | Freeze in the runner, not the prompt |
| Empty-search then blast-create | Invented entities | Empty result counts; create is gated |
What you do instead:
- Mark write tools in the catalog. They get the tightest fingerprint and tool caps.
- Forbid create-on-404 unless a named state allows it and a human has already approved the id class.
- Dry-run the write path in week one: “would write X.” Dual-write to a shadow field before hard writes.
- After any block, assert zero write executes in the rest of the run.
- Keep an unwind note in the escalate package: which row, which field, which timestamp.
- Write tools have a side-effect class in the catalog
- Create is not the default recovery from not-found
- Post-block write count is a monitored zero
- Someone can undo the last agent write from the escalate ticket
A prompt that says “be careful with writes” is not a brake. The runner refusing create_contact after get_contact 404’d is a brake.
How do I read a trace when the agent will not stop?
If you cannot see the repeated fingerprint, you will debug the model. Debug the execute list first.
| Question | Where to look | If you cannot answer |
|---|---|---|
| Which tool fired? | Execute spans / log lines, not chat prose | You do not have traces |
| Same args or new args? | Normalized arg hash, or a redacted args panel | You cannot fingerprint yet |
| Error class? | Status code / error_code, not the model’s summary | You threw away HTTP |
| Did the harness refuse? | A blocked flag | You only log after success |
| Did anything wait? | Timestamps between executes | You cannot tell grind from backoff |
| Why did the run end? | Terminal reason code | You will not trend this next week |
Operator pass on a stuck run (ten minutes, one run):
- Export executes in order: time, tool, arg hash, status, retryable, blocked.
- Group by arg hash. Any group with count ≥ 2 on a permanent class is the incident.
- Group by tool name. Any tool past its cap is the backstop catching wander.
- Check the last reason code. If it is generic
erroror missing, fix the catalog before you fix the prompt. - Check whether a summarize step sits between attempt 1 and attempt 2. If yes, confirm the ledger survived.
| Metric to start this month | Why | Hedge |
|---|---|---|
| Share of runs with any fingerprint count ≥ 2 | Loop pressure | Your baseline, not a published industry rate |
Share terminated tool_auth_error / policy_blocked | Are you escalating or grinding | Needs a reason catalog |
Executes after blocked | Brake bypass | Should be zero |
| Cost of runs with fingerprint ≥ 2 vs the rest | Money on the floor | Unit cost on your tools |
Do not wait for a weekly CSV to learn you hammered a 404 overnight. One alert on “blocked-fingerprint rate after a deploy” is enough to start. The complementary detector post is where the catalog of reason codes is specified; this page is where you use the trace as an operator.
When is a workflow enough instead of an agent?
If the honest path is “look up this id; if it exists, patch these fields; if it 404s, stop,” you do not need a model choosing tools in a loop. You need a workflow with an error branch.
Anthropic’s building-effective-agents essay draws the split: workflows follow predefined code paths; agents let the model direct process and tool use. They tell teams to find the simplest solution and add complexity only when it demonstrably helps. A same-failed-tool loop is often proof you added a director to a job that needed a branch.
| Job shape | Prefer | Why |
|---|---|---|
| Known id → known write → stop on 404 | Workflow | The miss is a branch, not a plan |
| Auth failure → page secrets | Workflow / runner | No model step adds a valid token |
| Policy refuse → stop | Runner gate | The model should not see a workaround |
| Ambiguous record match across systems | Agent, with caps | You need a search strategy — still capped |
| Open research with a write at the end | Hybrid: agent reads, workflow writes | Do not let research thrash a write tool |
Decision list:
- Can you draw the happy path and the 404 path without a model? Ship a workflow.
- Does the job need a search strategy under uncertainty? An agent may earn the loop — with fingerprint, tool cap, and escalate.
- Is the failure class auth or policy? Neither topology should retry. The runner stops.
- Are you using an agent because the demo looked smarter? Read the production-control spoke before you widen autonomy.
Guardrails you need either way:
- Typed tool schemas
- Tenant bound at the runner
- Fingerprint + same-tool cap on every tool the job can call
- Auth/policy escalate with an owner
- A terminal other than “the model said done”
An agent is not a retry engine. If retries are the product, you wanted a queue with backoff.
What should I not do while it is still looping?
Add “please don’t retry failed tools” to the system prompt as the only fix. It will drift on the next “be more persistent” edit.
Retry every error three times. Auth, policy, validation, and not-found are not transient. HTTP already sorted this. Your catalog should too.
Raise temperature, swap models, or buy a larger context window. Irrelevant to a deterministic 404 with the same args.
Log the loop and keep executing. Observation without a refuse path is a museum exhibit.
Raise the same-tool cap because the job is “hard.” Hard jobs escalate with evidence.
Treat an empty search as progress. [] on the same query is a hit.
Let the model pick a second write tool after the first 404. That is a workaround, not recovery.
| Temptation | Why it fails | Do this instead |
|---|---|---|
| Longer prompt | Drift + summarization | Fingerprint gate |
| Retry-everything | Permanent classes hammer | Classify; fail closed on unknown |
| New model | Same missing brake | Same harness |
| Unlimited turns | Bill with no artifact | Finite cap + escalate |
Hide 403 as 404 and keep searching | You deleted the escalate signal | Keep the status; page roles |
| “Just this once” disable the cap in prod | The exception becomes the path | Change the job type’s numbers in config, in staging first |
OWASP’s unbounded-consumption entry is the cost framing: uncontrolled inference (and, for you, uncontrolled tool executes) is how a pay-per-use system becomes a surprise invoice. Caps are not pessimism. They are how the job stays a job.
How do I measure this in production?
You asked how to evaluate it. Do not wait for a vendor “agent reliability” score. Count repeats.
| Question | Metric | Pass shape (set yours) |
|---|---|---|
| Is the identical hammer dying? | Fingerprint count ≥ 2 on permanent classes | Falls after the gate ships; never a claimed industry % |
| Is neighbor-id thrash dying? | Same-tool executes per run | Median at or under the cap |
| Are we escalating auth? | Runs with 401/403 that still executed again | Zero |
| Are we escalating policy? | Policy refuses that still hit the vendor | Zero |
| Did the brake hold after a prompt edit? | Golden case for a known bad id | Still red on second execute |
| Are writes safe? | Write executes after blocked | Zero |
Procedure for a baseline this week — no invented loop rate, just your numbers:
- Take the last 20 failed runs (or all of last week if you have fewer).
- For each, mark: identical-fingerprint repeat, same-tool wander, auth retry, policy retry, legitimate backoff, other.
- Write the counts on one page. That page is the baseline.
- Ship fingerprint + tool cap + auth/policy terminate.
- Repeat the 20-run sample after the change. Compare counts, not adjectives.
| You will be tempted to report | Report this instead |
|---|---|
| “Loops are down a lot” | Identical-fingerprint repeats: 11/20 → 2/20 on this sample |
| “The new model is better” | Same sample, same harness, one variable |
| “Zero loops” | Zero identical repeats; wander still has a tool-cap metric |
| A percent copied from a vendor blog | Your sample, dated, with the definition you used |
Golden-case the known offender so the measurement is not only a retro:
- Stub the permanent error
- Assert the second identical execute never hits the client
- Assert auth/policy never hits the client twice
- Assert the same-tool cap fires on a neighbor walk
- Force a summarize between attempts; assert the block still holds
If a prompt edit that “makes the agent more persistent” does not turn that case red, you are still measuring vibes.
What can I ship in a week?
You asked what to skip if you only have a week. Skip the dashboard redesign, the multi-agent oscillation work, and the prompt rewrite. Keep the refuse path.
Day-by-day, one job, write tools first:
- Day 1. Dump executes from five real failures. Classify error classes by hand. Name the identical hammer vs the wander.
- Day 2. Normalize + fingerprint for those tools. Unit test “same id, different key order.”
- Day 3. Pre-execute refuse on fingerprint. Auth/policy terminate on first hit. Unknown fails closed.
- Day 4. Same-tool cap on the write tool. Reason code on the run record. Escalate package template.
- Day 5. Golden case + a 20-run recount. If writes still execute after block, that is the only bug you work.
| If you only have two days | Do | Skip |
|---|---|---|
| Do first | Fingerprint refuse + auth/policy terminate | Pretty traces |
| Next | Same-tool cap on writes | Alternate-tool graphs |
| Proof | One golden bad-id case | Fleet-wide percentiles |
Checklist for the PR that actually ships:
- Fingerprint computed pre-execute
- Same-tool counter on the run
-
401/403/ policy refuse never retry - Reason code is an enum
- Escalate names an owner
- Golden case green
- Prompt text is not the brake
If you only have a day: fingerprint the worst write tool, cap it at 2, terminate auth. Coarse beats a better paragraph.
What does an agentic pilot install for this?
In a Spurlock Studios $1,500 · 5-day agentic pilot, same-failed-tool brakes are part of the thin harness for one job: fingerprints on write-capable tools, a same-tool cap, auth/policy terminate, a reason code, and at least one golden bad-id case. Full fleet dashboards and multi-agent cycle detection can wait. The identical hammer should not.
You leave with counts on your last sample, not a promised loop-rate. Architecture context stays in the operating manual. Start from /agentic.
| Pilot includes | Pilot does not include |
|---|---|
| One job, real tools, real error classes | A claim that loops cannot happen |
| Fingerprint + same-tool cap + escalate | A new model as the fix |
| Golden case in CI | Prompt-only “don’t retry” language as the control |
| Reason codes you can trend later | A full observability rebuild |
If the job is still a known branch, the pilot may tell you to ship a workflow. That is a successful week.
FAQ
Why does my agent keep calling the same failed tool?
Because the runner treats every model turn as progress and never fingerprints the retry or caps same-tool executes. The model sees an error and proposes the same call; nothing refuses the next execute. Auth and policy failures need an escalate path on the first hit, not another attempt with the same credentials.
How do I measure whether a same-tool retry cap is working?
Count identical-fingerprint repeats and same-tool executes on a dated sample of failed runs, then recount after the gate ships. Auth and policy classes should show zero second executes. Do not publish a vendor-style loop rate — report your sample, your definition, and whether the golden bad-id case still blocks.
What usually fails first when teams try this?
Error classification. Teams retry 401, 403, and policy refuses as if they were timeouts, or they fingerprint after the HTTP client so the bill already happened. The second miss is skipping the same-tool cap, so a 404 on one id becomes a neighbor walk that still looks like “progress” in the trace.
How long does this take to show results?
A coarse fingerprint plus a write-tool cap plus auth terminate can show up on the next failing run — often inside a week if you already have execute logs. Trendable reason codes take longer because you need a catalog and a sample. Do not wait on a dashboard to refuse the second identical call.
What should I skip if I only have a week?
Skip prompt rewrites, model swaps, multi-agent oscillation work, and a full observability rebuild. Fingerprint the worst write tool, cap that tool, terminate auth and policy on first hit, and add one golden bad-id case. That is the refuse path. Everything else is furniture.
When is this not worth doing yet?
If the job is a known lookup-then-stop path, do not add an agent so you can then cap it — ship a workflow with an error branch. If you cannot list the tools or see executes, instrument one run before you write a fingerprint function. If writes are already unconstrained in production, freeze writes first; then add retry brakes.
CTA
Cap the retry. Escalate the auth error. Then decide if the job still needs an agent.
What questions does this article answer?
- Why does my agent keep calling the same failed tool?
- Because the runner treats every model turn as progress and never fingerprints the retry or caps same-tool executes. The model sees an error and proposes the same call; nothing refuses the next execute. Auth and policy failures need an escalate path on the first hit, not another attempt with the same credentials.
- How do I measure whether a same-tool retry cap is working?
- Count identical-fingerprint repeats and same-tool executes on a dated sample of failed runs, then recount after the gate ships. Auth and policy classes should show zero second executes. Do not publish a vendor-style loop rate — report your sample, your definition, and whether the golden bad-id case still blocks.
- What usually fails first when teams try this?
- Error classification. Teams retry `401`, `403`, and policy refuses as if they were timeouts, or they fingerprint after the HTTP client so the bill already happened. The second miss is skipping the same-tool cap, so a 404 on one id becomes a neighbor walk that still looks like “progress” in the trace.
- How long does this take to show results?
- A coarse fingerprint plus a write-tool cap plus auth terminate can show up on the next failing run — often inside a week if you already have execute logs. Trendable reason codes take longer because you need a catalog and a sample. Do not wait on a dashboard to refuse the second identical call.
- What should I skip if I only have a week?
- Skip prompt rewrites, model swaps, multi-agent oscillation work, and a full observability rebuild. Fingerprint the worst write tool, cap that tool, terminate auth and policy on first hit, and add one golden bad-id case. That is the refuse path. Everything else is furniture.
- When is this not worth doing yet?
- If the job is a known lookup-then-stop path, do not add an agent so you can then cap it — ship a workflow with an error branch. If you cannot list the tools or see executes, instrument one run before you write a fingerprint function. If writes are already unconstrained in production, freeze writes first; then add retry brakes.
Last reviewed
AI Agents
AI Agents Budtender FAQ that will not invent a strain benefit
A floor FAQ agent answers hours, pickup rules, and SKUs from approved copy — then hard-stops before inventing a medical claim or a COA.
AI Agents Who is accountable when an agent acts (refunds, emails, writes)
A named human owns every agent refund, email, and write. Policy gates sit before irreversible tools; the model is not a person and cannot absorb the blame.
AI Agents Why Pass Rate Lies: Revision Rate, Trajectories, and Coverage
Pass rate flatters bad agents. Gate deploys on revision rate, trajectory scores, eval coverage, and cost per successful task—not a single green percentage.
AI Agents Agentic Systems: An Operating Manual for Multi-Agent Work That Ships
An agentic system is evaluators, policy gates, sandboxes, and kill switches — not a chat window — so production tool work survives contact with real data.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.