Spurlock Studios
Contact
Share LinkedIn X
A scuffed work smartphone with a blank glowing circular button. Thesis: AGENTS LOOP FAILED TOOLS PROGRESS.

Your agent keeps calling the same failed tool because the harness treats every model turn as progress. The model sees an error, believes another attempt will help, and you never fingerprint the call as a no-progress repeat. Longer prompts do not fix a missing loop detector.

This spoke sits under the Agentic Systems Operating Manual. It owns detection: fingerprint the call, honor retryable: false, cap turns, terminate with a reason. Traces and the ops screen live on observability for agents. Start a thin harness from /agentic.

The short answer

  • A legitimate retry changes something material: args, backoff, or a different tool. A no-progress loop repeats the same fingerprint.
  • Tool-error channels (isError, is_error) feed the model. They do not refuse the next execute. That refusal is harness work.
  • Context summarization often re-triggers loops by dropping the “this already failed” evidence. Persist a failure ledger outside the prompt.
  • Fix it in the harness: fingerprint tool calls, honor retryable: false, cap turns, terminate with a reason code, alert on loop-rate spikes.
  • Prompting “don’t retry forever” is a hint, not a control. Golden-case the loop once you kill it — or it returns after the next prompt edit.

What counts as a no-progress loop?

Ops needs a crisp distinction. Without it, every retry looks like diligence and every stop looks like you gave up.

SignalLegitimate retryNo-progress loop
Tool nameSame or a real alternateSame
ArgumentsChanged (id, page, filter)Identical or equivalent after normalize
TimingBackoff / jitter, often Retry-AfterImmediate hammer
Prior resultTransient (429, timeout, some 5xx)Permanent (404, auth, validation)
Harness viewNew fingerprint or marked retryableSame fingerprint ≥ N times

Rule of thumb: if a human watching the trace would say “it’s doing the same thing again,” the harness should already have stopped it.

  • Same tool name
  • Same normalized args
  • Same error class, or empty result treated as “try again”
  • No backoff
  • No new state checksum

Four of five is enough to refuse the execute. Do not wait for the token budget to notice.

Why do models retry even when the error is clear?

Models optimize for completing the user job. An error message is just more context. Unless the harness injects a hard stop, the next plan is often “try the tool again.” That is rational under incomplete control — and expensive under write tools or paid APIs.

After 20,000+ hours on agentic systems and 500+ automations, the pattern is boring: the model is not “confused.” The cage never told it the attempt was terminal.

Common fuel for loops:

  1. Vague tool errors — "failed" with no retryable flag
  2. Silent empty results — [] treated as “search again with the same query”
  3. Prompt pressure — “keep going until done” with no terminate authority
  4. Missing memory of failure — the ledger lived only in chat tokens
  5. Framework defaults that look like safety — a high step cap is not a job-type budget
FuelWhat the model seesWhat the harness should do
"failed"Maybe a blipClassify; default unknown → not retryable
[]Maybe a bad queryCount empty-same-query as a fingerprint
“Don’t stop”Permission to grindIgnore; enforce caps anyway
Summarized historyA fresh planInject the failure ledger every turn
max_turns=NoneInfinite roomRefuse that default in prod

Do not moralize the model. Instrument the harness.

Why don’t longer prompts stop the loop?

A system prompt that says “if a tool fails with not-found, stop” is a suggestion. The next deploy that “improves helpfulness” will soften it. The next summarizer will drop it. The next model turn will rediscover the original plan.

ControlLives whereSurvives a prompt edit?Survives summarization?
“Please don’t loop”System promptNoNo
Few-shot of a stopped runChat tokensUntil the next rewriteUntil compress
Fingerprint tableHarness storeYesYes
retryable: false gatePre-executeYesYes
max_turns + no-progress capRunner configYesYes
Terminal reason codeRun recordYesYes

The take: no-progress detection beats longer prompts because detection is not text. It is a pre-execute check against a ledger the model cannot overwrite.

If you only add language, you bought a temporary mood. If you add a fingerprint gate, the mood can change and the brake still fires.

How do you fingerprint tool calls in the harness?

Fingerprinting is the core no-progress signal. Compute it before execute, not after the bill.

fingerprint = hash(tool_name + normalize(args) + side_effect_class)

Normalization rules matter. Two JSON objects that a human would call “the same call” must hash the same.

RuleDoDo not
Object keysSortHash insertion order
Volatile fieldsStrip request_id, injected timestampsInclude harness-owned noise
IdsTrim; lowercase only where the system is case-insensitiveLowercase a case-sensitive SKU
AuthExclude headers and tokensHash the bearer token
Empty vs missingCanonicalize null / omitted per tool schemaTreat {} and {filter:null} as different without a rule
Side-effect classInclude read / write / side_effectHash reads and writes as if they were the same risk

Store per run:

FieldPurpose
fingerprintIdentity of the attempt
countHow many times this run hit it
first_error_code / last_error_codeStability of failure
retryableFrom the tool or a classifier
blockedHarness refused further executes
last_atFor backoff and alerts

Policy example (illustrative defaults — tune per job type):

  1. Same fingerprint + retryable: false → refuse immediately, reason tool_no_progress
  2. Same fingerprint + retryable → allow up to 2 retries with backoff, then refuse (tool_retry_exhausted)
  3. Distinct fingerprints that still share tool + error class → soft warn; escalate after a threshold
  4. Empty result + same fingerprint → count it. “Nothing found” is a result, not a license to hammer

The model can propose the call. The harness decides whether it runs.

  • Fingerprint computed pre-execute
  • Normalize function has a unit test per tool
  • Volatile fields listed and excluded
  • Block path never reaches the HTTP client

What must retryable:false mean?

Tools are part of the control loop. An error string the model can reread is not a policy. The payload has to be machine-readable, and the harness has to obey it.

{
  "ok": false,
  "error_code": "contact_not_found",
  "retryable": false,
  "message": "No contact for id=…"
}

This is the same idea Temporal documents for activities: mark a failure non-retryable at the throw site, or list types the runtime must never retry — invalid input, missing records, authorization — because repeating the same call will never succeed (Temporal: non-retryable errors). Their TypeScript RetryPolicy matches on error type names, not message strings. Your agent tools need the same discipline: a flag and a code, not a paragraph.

Error classretryableHarness action
Not found / validationfalseBlock fingerprint; maybe one alternate tool
Auth / permissionfalseTerminate tool_auth_error
Rate limit / timeouttrueBackoff, capped retries
5xx / upstream bliptrueBackoff, then escalate
Unknowndefault false in prodFail closed

Defaulting unknown errors to retryable is how you buy infinite loops. Prefer fail closed; loosen per tool after evidence.

HTTP already sorted this for you. RFC 9110 defines Retry-After as the wait before a follow-up. MDN’s Retry-After notes the common pairings: 503 (temporary unavailability) and 429 (rate limit). Those are retryable with a clock. 404, 401, and 422 are not a clock. They are a stop.

HTTP / tool signalRetryable?Clock
429 + Retry-AfterYes, cappedHonor the header
503 + Retry-AfterYes, cappedHonor the header
Timeout / connect resetYes, cappedJittered backoff
404 / not foundNoNone
401 / 403NoNone — fix creds
422 / schema rejectNoNone — fix args

If your tool wraps HTTP and throws away the status, you deleted the only honest retryable bit. Keep it.

Why do isError and is_error feed the loop?

Most tool protocols are designed so the model can recover. That is useful for a typo in a search query. It is lethal for a permanent not_found if the harness treats every error as “more context.”

The Model Context Protocol tools spec (2025-06-18) splits two channels. Protocol errors (unknown tool, invalid arguments) are JSON-RPC errors. Tool execution errors — API failures, bad input, business-rule rejects — come back as a successful result with isError: true so the model can read them. Anthropic’s Messages primer does the same on the client side: if the tool throws, return a tool_result with "is_error": true. Their stop-reason guide is equally clear that tool_use means “run the tool and continue,” not “the job is done.”

ChannelWho sees itWho can refuse the next call
MCP JSON-RPC errorClient / hostHost, if it maps the code
MCP isError: trueModelHarness, if it fingerprints
Anthropic is_error: trueModelHarness, if it fingerprints
Anthropic stop_reason: tool_useYour loopYour loop — this is “continue”
Your retryable: falseHarnessHarness — this is “stop”

The failure mode: you faithfully return isError: true / is_error: true, the model “helpfully” retries the same args, and you execute it because the protocol said the result was for the model. The protocol is not the brake. You are.

  • Map protocol-invalid-args to retryable: false before the model turn
  • Map execution isError through the same classifier as HTTP
  • Never treat “the model asked again” as permission to run again

Why cap turns and no-progress events?

Infinite max_turns is a production bug, not a feature. A high framework default is also not a job-type budget. You need two ceilings: steps, and repeated fingerprints.

The OpenAI Agents SDK runner is explicit: if the run exceeds max_turns, it raises MaxTurnsExceeded. A turn is one model invocation, including the tool calls in that step. Pass max_turns=None and you disable the limit. That is the dangerously infinite default — not a number, an off switch. Their error_handlers["max_turns"] can synthesize a final output. Use it to emit a reason code, not to pretend the job finished.

LangGraph’s graph API caps super-steps with recursion_limit and raises GraphRecursionError when you blow it. As of version 1.0.6 the documented default is 1000 steps. That is a hang-preventer, not a CRM-update budget. They also expose RemainingSteps so you can route to a completion node before the exception. Prefer the proactive path: terminate with max_turns while you still have a ledger to attach.

CrewAI documents max_iter as the iteration cap that forces a best-effort answer. The learn page lists a default of 25. The edge concepts page lists 20, plus max_retry_limit default 2. As of August 2026 those surfaces disagree on the default — treat the existence of the cap as the lesson, and set the number per job type yourself.

CeilingWhat it stopsWhat it misses
max_turns / recursion_limit / max_iterEndless step countSame fingerprint on step 3 of 12
No-progress fingerprint capThe 404 hammerA wander across new bad ids
Wall-clock / token budgetSpendA cheap tight loop that still writes
RemainingSteps / escalate nodeUgly exceptionsNothing, if you never route to it

Recommended pairing per job type (illustrative — measure your successful trajectories):

Job typeTurn capSame-fingerprint capNotes
Read-only lookup6–101 if not retryable; 3 if retryableEmpty result counts
Single CRM write8–121 on retryable: falseWrites after block = bug
Multi-step research20–402 per fingerprintStill not 1000
Human-in-the-loopUntil approve or budget1Escalate, don’t grind

Cap turns and cap no-progress events. One without the other still burns money.

Why terminate with a reason code, not a paragraph?

A stop that ops cannot group is a stop you will not trend. Free-text “the agent got stuck” is how loop-rate spikes hide in a chat export.

Recommended terminal reason codes for this failure family:

CodeWhenWhat to attach
tool_no_progressFingerprint blocked after policyLedger row, tool, arg hash, error code
tool_retry_exhaustedRetryable path used upAttempt count, backoff log
max_turnsStep budget hitTurn count vs cap, last fingerprints
tool_auth_errorPermission / credsRedacted scope, not the secret
escalateHuman pathFailure ledger + last artifact

Do not overload generic error. You cannot alert on mush.

OpenAI’s runner already treats max-turns as a first-class kind you can handle. Mirror that: your harness should have a catalog, a single writer, and a field on the root span. Observability for agents is where that field becomes a dashboard tile. This post is where you decide the code before the tile.

  • Reason code is an enum, not a sentence
  • Every terminate path sets one
  • Ledger travels with escalate
  • Prompt text is not the reason

Why does summarization re-trigger loops?

Long runs compress history. Summarizers keep “goals” and drop “we already called crm.get_contact with id X and got not_found three times.” The model, seeing a fresh window, rediscovers the same plan.

CrewAI’s own agent attributes include respect_context_window (default true): keep messages under the window by summarizing (edge concepts). That is the right instinct for tokens and the wrong place to store terminal failures. If your only copy of retryable: false lives in a tool_result the summarizer will squash, you will loop after the compress.

Controls that survive summarization:

ControlWhere it livesInjected how
Failure ledgerStore beside tracesStructured system state every turn
Blocked fingerprint setSame storePre-execute deny list
Last N failure fingerprintsVerbatim, not summarizedAppend-only block
Terminal error for this runNever eligible for compressPin until the run ends

Procedure when you add a summarizer:

  1. Persist (tool, arg_hash, error_code, count, last_at, retryable) outside the prompt.
  2. Inject that ledger as structured state, not chat prose.
  3. Mark terminal tool errors for the current run as not-summarizable.
  4. On summarize, keep the last N distinct failure fingerprints verbatim.
  5. Golden-case: force a compress mid-run and assert the second identical call is still refused.

If failure evidence only lives in chat tokens, compression will resurrect the loop.

What is a legitimate retry vs the 404 hammer?

Illustrative — not a client result. Agent is told to update contact c_1842. Tool returns 404 contact_not_found, retryable: false. Without fingerprinting, the agent retries the same id twelve times, then invents nearby ids. With fingerprinting: first failure records the fingerprint; second proposal is refused; run terminates tool_no_progress with an escalate package so a human can verify the id.

Cost difference is not subtle when the tool is a paid enrichment API. It is also not subtle when the tool is a write: twelve “update” attempts on a missing row are twelve chances to create junk if the API upserts.

StepNo detectorWith detector
1Execute, 404Execute, 404, ledger row
2Execute againRefuse, tool_no_progress
3–12Execute againNever reached
13Invented idsHuman sees the real id problem
Bill12× tool + 12× model1× tool + 2× model

A legitimate retry looks different on the same table: 429, Retry-After: 8, same fingerprint, retryable: true, wait, one more execute, success. That is diligence. The harness allowed it because the flag and the clock said so — not because the model asked nicely.

ABAB handoff oscillations are a second species. Agent A hands to Agent B, B hands back to A, neither advances the artifact. Same-tool loops repeat one fingerprint. Oscillations need a package hash and a cycle check, not a longer tool-name allowlist. Do not stretch tool fingerprinting to cover them.

PatternDetectionFix
Same-tool loopFingerprint countBlock tool; reason code
ABAB handoffHandoff graph cycle / identical package hashBreak cycle; merge agents or escalate
Alternate-tool thrashTool set cycles without state changeRequire a state checksum to move

State machines own legal states and revision ceilings. Harness guards own the refuse-to-execute decision inside act. If you only have a state machine, you can still spin inside act. If you only have fingerprinting, you can still wander illegal states. Ship both. The cage is in the operating manual; this spoke is the detector.

What must traces show for a loop?

If you cannot see the repeated fingerprint, you will debug the model. OpenTelemetry’s 2026 GenAI observability write-up is the right picture: one invoke_agent tree with child chat spans and execute_tool spans for each invocation. Their conventions also warn that tool arguments are opt-in because they hold secrets. You need the hash by default, the redacted args when a human opens the run.

Span / fieldLoop question it answers
execute_tool + gen_ai.tool.nameWhich tool
Arg hash / fingerprint attributeSame call or new call
error.type / your error_codePermanent vs transient
retryableDid the tool tell the truth
blockedDid the harness refuse
Terminal reasonWhy the run ended
Cost on the rootWhat the loop spent

Alert when blocked-fingerprint rate exceeds a trailing baseline after a deploy or prompt change. That is how you catch a “helpful” edit that removed the stop language — or, better, prove the harness never depended on that language.

MetricWhy
% runs with any blocked fingerprintLoop pressure
avg duplicate fingerprints per runSeverity
runs terminated tool_no_progressHard stops working
cost of runs with loop≥1Money on the floor

The dashboard that shows those tiles is observability for agents. Do not wait for a weekly export to learn you hammered a 404.

Idempotency is a cousin, not a substitute. Stripe’s idempotent requests let a client retry a POST with the same Idempotency-Key so a network blip does not create two objects. That protects the server from duplicate side effects. It does not tell your agent to stop proposing a call that already returned a permanent error. Use idempotency keys on write tools and refuse the no-progress fingerprint. One without the other is how you get a safe duplicate and an infinite bill.

How do you add no-progress detection this week?

  1. List write-capable tools first. Reads can loop too; writes are the ones that leave a mess.
  2. Define fingerprint normalization for each tool. Write the unit tests before the gate.
  3. Add a per-run failure ledger beside traces: fingerprint, count, codes, retryable, blocked.
  4. Enforce retryable from tool payloads. Default unknown → false in production.
  5. Set max_turns (never None in prod) and a max blocked-fingerprint count per job type.
  6. Emit reason codes on terminate. Wire one alert on loop-rate spike.
  7. Add a golden case that expects tool_no_progress for a known bad id.
  8. Confirm summarization preserves the ledger. Force a compress in the test.

Checklist for the PR:

  • Fingerprint computed pre-execute
  • Block path refuses model retries
  • retryable: false never hits the network twice
  • Reason code visible in the ops dashboard
  • Golden case green
  • Alert stubbed (even if the threshold is temporary)
  • max_turns=None / unlimited recursion is a lint failure

If you only have a day: steps 2, 4, 5, and 7. A coarse hash and a finite cap beat a better prompt.

How do you golden-case a known loop?

Capture a production offender once. Prompts will change. The case keeps the detector honest.

  1. Freeze tool stubs that return the permanent error (404, retryable: false).
  2. Assert the agent proposes the tool (optional — useful while you still trust the prompt).
  3. Assert the harness blocks the second identical fingerprint.
  4. Assert the terminal reason is tool_no_progress (or your chosen code).
  5. Assert no write tools ran after the block.
  6. Force a context summarize between attempt 1 and attempt 2. Assert the block still holds.
  7. Repeat with a retryable 429 and assert one backoff retry is allowed, then tool_retry_exhausted.
AssertionPasses whenFails when
Second identical executeHarness refusesHTTP client is called
Reason codetool_no_progressGeneric error
Writes after blockZeroAny
After summarizeStill refusedLoop returns
Retryable path≤ N executesImmediate refuse or unbounded

Keep that case in CI. A prompt edit that “makes the agent more persistent” should go red.

What should you not do to stop agent loops?

Raise temperature and hope. Irrelevant to deterministic 404s.

Add “please don’t loop” to the system prompt as the only fix. It will drift.

Retry all errors three times. Auth and validation errors are not transient. Temporal’s non-retryable list exists because this lesson is older than agents.

Log loops without terminating. Observation without a brake is a museum exhibit.

Set max_turns=None or raise recursion_limit to 1000 “for hard tasks.” Hard tasks need escalate, not eternity. LangGraph’s 1000-step default is a last-resort hang guard, not a product setting.

Treat Stripe-style idempotency as loop detection. Safe duplicates are not a stop.

Count empty search results as progress. [] on the same query is a fingerprint.

TemptationWhy it failsDo this instead
Longer promptDrift + summarizationFingerprint gate
Retry-everythingPermanent errors hammerretryable flag
Unlimited turnsBill with no artifactFinite cap + escalate
Trace-onlyYou watch the fireRefuse the execute
New modelSame missing brakeSame harness

What is the pilot minimum for loop detection?

In a Spurlock Studios $1,500 · 5-day agentic pilot, no-progress detection is part of the thin harness: fingerprinting on write-capable tools, turn caps, reason codes, and at least one golden loop case. Full multi-agent oscillation detection can wait; same-tool loops should not.

Start from /agentic. Architecture context: operating manual. The screen that shows the reason code: observability for agents.

FAQ

What max_turns default is dangerously infinite?

Any default that is null, max_turns=None, zero-means-unlimited, or set in the thousands “just in case.” The OpenAI Agents SDK will disable the limit if you pass None; LangGraph’s documented default of 1000 super-steps (as of 1.0.6) is a hang-preventer, not a job budget. Pick a finite cap per job type that matches real successful trajectories, then escalate — do not let the model grind until the budget burns.

Should tools return retryable: false?

Yes for permanent failures: not found, validation, auth, and business-rule rejects. Transient classes (rate limit, timeout, some 5xx) return true with harness-enforced caps and, when the server sends it, Retry-After. Unknown errors should default to non-retryable in production. Temporal’s non-retryable activity errors are the same rule with different nouns.

How do ABAB handoff oscillations differ from same-tool loops?

Same-tool loops repeat one fingerprint. ABAB oscillations bounce work between agents without artifact progress. Detect them with handoff-package hashes and cycle checks, not only tool fingerprints. Do not stretch a tool-name allowlist to cover a graph cycle.

Where do state machines help vs harness guards?

State machines own legal states, transitions, and revision ceilings. Harness guards own fingerprinting, retryable policy, and refusing duplicate executes inside act. You need the cage and the detector. A legal act state that re-runs the same 404 is still a loop.

What reason code should terminate the run?

Use a dedicated code such as tool_no_progress when a fingerprint is blocked, and tool_retry_exhausted when retryable attempts are spent. Use max_turns when the step budget hits. Do not overload generic error — ops cannot trend mush, and you cannot alert on a paragraph.

How do I add a golden case for a known loop?

Stub the permanent tool error, assert the harness blocks the repeated fingerprint, assert the terminal reason code, and assert no further writes. Add a second case that forces summarization between attempts. Keep both in CI so prompt edits cannot delete the brake.

CTA

Stop paying for the same failed tool call.

/agentic · /contact?intent=agentic-pilot

FAQ

What questions does this article answer?

What max_turns default is dangerously infinite?
Any default that is null, `max_turns=None`, zero-means-unlimited, or set in the thousands “just in case.” The OpenAI Agents SDK will disable the limit if you pass `None`; LangGraph’s documented default of 1000 super-steps (as of 1.0.6) is a hang-preventer, not a job budget. Pick a finite cap per job type that matches real successful trajectories, then escalate — do not let the model grind until the budget burns.
Should tools return retryable: false?
Yes for permanent failures: not found, validation, auth, and business-rule rejects. Transient classes (rate limit, timeout, some 5xx) return `true` with harness-enforced caps and, when the server sends it, `Retry-After`. Unknown errors should default to non-retryable in production. Temporal’s non-retryable activity errors are the same rule with different nouns.
How do ABAB handoff oscillations differ from same-tool loops?
Same-tool loops repeat one fingerprint. ABAB oscillations bounce work between agents without artifact progress. Detect them with handoff-package hashes and cycle checks, not only tool fingerprints. Do not stretch a tool-name allowlist to cover a graph cycle.
Where do state machines help vs harness guards?
State machines own legal states, transitions, and revision ceilings. Harness guards own fingerprinting, `retryable` policy, and refusing duplicate executes inside `act`. You need the cage and the detector. A legal `act` state that re-runs the same 404 is still a loop.
What reason code should terminate the run?
Use a dedicated code such as `tool_no_progress` when a fingerprint is blocked, and `tool_retry_exhausted` when retryable attempts are spent. Use `max_turns` when the step budget hits. Do not overload generic `error` — ops cannot trend mush, and you cannot alert on a paragraph.
How do I add a golden case for a known loop?
Stub the permanent tool error, assert the harness blocks the repeated fingerprint, assert the terminal reason code, and assert no further writes. Add a second case that forces summarization between attempts. Keep both in CI so prompt edits cannot delete the brake.
Sources

Last reviewed

More from this lane

AI Agents

All →
Start a pilot