Dead Letter Queues for Automations: Where Failed Work Goes to Be Fixed
Park failed n8n work in a dead-letter queue with a replay path. Unbounded retries on a bad payload become vendor rate-limit incidents instead of recovery.
William Spurlock Founder — Spurlock Studios Updated 20 MIN
Retries feel responsible. Unbounded retries on a bad payload are how you turn one failure into a rate-limit incident.
A dead-letter queue (DLQ) is the opposite instinct: when work cannot complete safely, park it with enough context for a human to fix and replay. That is production discipline for n8n and every other automation rail. Across 500+ automations, the graphs we still trust are the ones that stop hitting the vendor after a budget, write a row a human can open, and resume from the failed step instead of inventing a second write.
The spine lives in the Production n8n handbook. The wake-up path is a readable n8n error workflow. This spoke is the parking lot and the replay handle.
The short answer
- Retry symptoms of the network. Dead-letter symptoms of the data or the design.
- Bound every retry. n8n’s own node settings describe Retry On Fail as rerunning “until it succeeds.” That sentence is the incident.
- Park the original input, the error, the execution id, and an owner. The Error Trigger does not give you the payload by itself.
- Replay from the failed step with an idempotency key. Full-graph replay after a partial write is how you get duplicate CRM rows.
- Alert and store are different jobs. Slack pulls a human to the queue. The queue is a table with status.
What is a dead-letter queue in an automation graph?
In messaging systems, a DLQ holds messages a consumer could not process. Amazon SQS moves a message after maxReceiveCount. RabbitMQ republishes to a dead-letter exchange on reject, TTL, overflow, or delivery-limit. Azure Service Bus keeps a $deadletterqueue subqueue on every entity. Pub/Sub forwards to a dead-letter topic after a configured attempt budget.
An n8n graph is not a broker. The contract is the same:
| Contract piece | Messaging system | Automation graph |
|---|---|---|
| Isolate failed work | Separate queue / topic / subqueue | Table, Airtable base, or Redis stream |
| Bound attempts first | maxReceiveCount, delivery-limit, max delivery attempts | Node Retry On Fail + Wait, then stop |
| Keep the original body | Message payload + death headers | Redacted rawInput JSON |
| Explain why it parked | DeadLetterReason, x-death, redrive metadata | errorClass + failed node + message |
| Give a replay path | Redrive / resubmit / new publish | Resume node or idempotent upsert |
| Name an owner | Alarm on DLQ depth | owner on the row, not “engineering” |
The medium is optional. The contract is not. If you cannot point at a row, an owner, and a replay path, you have logs — not a DLQ.
- Original input stored (redacted)
- Error and failed node stored
- Execution identity stored
- Owner assigned from workflow metadata
- Status is
openuntil replayed or discarded - Replay path documented for that workflow
Why do unbounded retries become rate-limit incidents?
One poison payload plus “keep trying” is a load test against someone else’s API, billed to you, and often shared across every workflow that uses the same credential.
n8n is explicit about the failure mode. When a node hits a rate limit it errors. If the service returns HTTP 429, the node output is The service is receiving too many requests from you — documented on the rate-limit page. The same page’s fix is a pause, not an infinite loop: set Wait Between Tries (ms) above the vendor window, or batch with Loop Over Items and Wait.
The trap is the default wording on Retry On Fail: the node reruns until it succeeds. A schema failure never succeeds. A revoked token never succeeds. An invalid email never succeeds. Those runs keep firing until the vendor 429s the whole workspace.
| Retry shape | What it does | What it becomes |
|---|---|---|
| No retry | First hard failure stops the graph | Safe for poison; noisy if the network blipped |
| Bounded + backoff | N tries, then fail the execution | Correct for 503 / timeout / 429 |
| Unbounded Retry On Fail | Rerun until success | One bad item → shared 429 → every flow stalls |
| Cron that re-pulls the same batch | Replays the whole window | Yesterday’s poison plus today’s volume |
| Error workflow that retries the source | Handler becomes a second hammer | Double the 429s, plus a muted Slack channel |
Failure mode we see: a HubSpot property rename lands on Friday. The nightly sync retries each of 400 contacts. By Saturday morning the credential is 429’d, the inbound lead webhook shares that credential, and sales thinks “the automation is down” — which is true, and self-inflicted.
- Classify the error before you retry.
- Retry only transient classes, with a ceiling and a wait.
- Fail the execution when the budget is gone.
- Let the error workflow park the item.
- Do not let the error workflow call the same vendor again.
When should you retry versus dead-letter?
Retry symptoms of the network. Dead-letter symptoms of the data or the design. Write the rule on the workflow, not in someone’s head at 11pm.
| Situation | Retry? | DLQ? | Why |
|---|---|---|---|
| HTTP 503 / timeout | Yes, bounded + backoff | After budget | Vendor may recover |
| HTTP 429 / rate limit | Yes, wait longer than the window | After budget | Retrying faster is the incident |
| Validation / schema failure | No | Immediately | Same payload will fail forever |
| Auth expired / 401 | No | Yes, and pause the workflow | Fix the credential first |
| Email API: invalid address | No | Yes, or a CRM hygiene path | Not a network problem |
| CRM create succeeded, later step failed | Not the whole flow | Yes, with partial-state note | Full retry duplicates |
| Unknown error | Once, maybe | Yes if it fails again | Do not invent a third class |
| Read-only enrichment miss | Optional | Only if the business needs the gap closed | Logging can be enough |
Decision list:
- Is the same payload expected to succeed if we wait? If no, DLQ now.
- Did an earlier node already write? If yes, DLQ with resume instructions — do not restart.
- Is the credential shared? If yes, pause before you retry a 401 across every graph.
- Is this customer-facing money or messaging? If yes, SEV1 on the DLQ row, not a digest.
If you cannot fill that list for a workflow, you are not ready to turn Retry On Fail on.
How do messaging systems already implement this?
Copy the broker contract. Do not invent a softer one because the rail is n8n.
Amazon SQS isolates unconsumed messages so you can inspect them and redrive later. The source queue’s redrive policy sets maxReceiveCount. A low count (1) parks on the first receive failure. AWS’s own guidance: set it high enough for real retries, not so high that poison sits in the hot path. For standard queues, message expiration stays tied to the original enqueue timestamp. If the item spent a day in the source queue and the DLQ retains four days, you have three days left. Set DLQ retention longer than the source queue, or the park evaporates.
RabbitMQ dead-letter exchanges republish when a consumer nacks without requeue, a TTL fires, a queue hits its length limit, or a quorum queue exceeds its delivery-limit. The broker records x-death (queue, reason, count). Reasons are rejected, expired, maxlen, delivery_limit. If the target exchange is missing, messages are silently dropped. A cycle with no rejection in the loop is dropped on purpose.
Azure Service Bus gives every queue and subscription a $deadletterqueue that you do not create and cannot delete. Default maximum delivery count is 10. After that, reason is MaxDeliveryCountExceeded. There is no automatic cleanup — messages stay until you complete them. Application code can dead-letter on purpose and should set DeadLetterReason plus a description a human can read.
Pub/Sub forwards to a dead-letter topic after an approximate attempt budget. Default maximum delivery attempts is 5, minimum 5, maximum 100. Counts are best-effort and can reset on inactive pull subscribers. You still need a subscription on the dead-letter topic, or you parked into a hole.
| System | Bound | Park target | Replay | Gotcha |
|---|---|---|---|---|
| SQS | maxReceiveCount | Separate queue | Redrive | Standard-queue TTL does not reset on move |
| RabbitMQ | reject / TTL / maxlen / delivery-limit | DLX → queue | Republish | Missing DLX drops the message |
| Service Bus | max delivery count (default 10) | $deadletterqueue | Resubmit | No auto-purge; you must complete |
| Pub/Sub | max delivery attempts (default 5) | Dead-letter topic | New subscription consume | Attempt count is approximate |
| n8n graph | You must invent the bound | You must invent the table | You must invent resume | Error Trigger is not the payload store |
n8n will not do this for you. The Error Trigger is the wake-up. The table is the queue.
How do you wire an n8n error workflow as the park path?
n8n’s error-workflow docs are short: for each workflow, set an error workflow in Workflow Settings. It runs if that execution fails. It must start with an Error Trigger. One handler can serve many graphs.
Procedure:
- Create
Error Handler — Productionwith an Error Trigger as the first node. - Normalize the payload (execution vs trigger-node shape).
- Classify
errorClass. - Write the DLQ row. If the write fails, that is a SEV1 of its own.
- Alert with a deep link and a next action. Details live in the error-workflow spoke.
- On every production graph: Options → Settings → Error workflow → Error Handler — Production → Save.
- Use Stop And Error when a validation node should fail the execution on purpose and hit the same handler.
Facts that bite:
- You cannot test the Error Trigger with Execute Workflow in the editor. It runs when an automatic workflow errors.
execution.idandexecution.urlrequire a saved execution. They are missing if the trigger node failed, because the workflow never executed.execution.retryOfis present only on a retry of a failed execution. Tag those so you do not open a second incident.- Continue / Continue (using error output) can swallow the error. From n8n’s point of view the execution handled it. The error workflow will not fire. Pair Continue with an explicit DLQ write on that branch, or do not use it on money paths.
- A workflow that contains an Error Trigger uses itself as its error workflow by default. Do not drop a test trigger onto a billing graph.
| Error Trigger field | Present when | DLQ use |
|---|---|---|
workflow.name / workflow.id | Almost always | Which graph |
execution.id / execution.url | Execution saved | Correlation + operator link |
execution.lastNodeExecuted | Mid-workflow failure | Where it died |
execution.error.message / stack | Mid-workflow failure | Why; trim the stack |
execution.retryOf | This run is a retry | Dedupe the incident |
trigger.error.* | Trigger-node failure | Fallback when execution{} is thin |
The Error Trigger does not include the input that caused the failure. Persist the payload before the risky write, or fetch execution details with an n8n API credential after. An alert without rawInput is a page with no work order.
If you save executions, the handler can call GET /executions/{id} and pull the failed node’s input out of run data. That fetch needs an n8n API credential on the error workflow, not on the graph that just died. A 401 on that fetch is an unhandled error inside the handler — keep it boring: timeout, one retry, then write the DLQ row with rawInput: unavailable rather than failing the park.
What payload must the DLQ row store?
If a human cannot replay from the row, the row is a log line. Minimum fields:
id
openedAt
workflowName
workflowId
executionId
executionUrl # or "unavailable — trigger failure"
failedNodeName
errorMessage # one line
errorStack # trimmed, optional
errorClass # transient_exhausted | schema | auth | partial | unknown
rawInput # redacted JSON
idempotencyKey # if any
businessId # customer, invoice, lead — what ops searches
partialState # e.g. hubspotContactId already created
owner
severity # SEV1 | SEV2 | SEV3
status # open | in_progress | replayed | discarded
assignee
notes
replayedAt
replayedBy
| Field | Why it earns a column | Failure if omitted |
|---|---|---|
rawInput | Replay source of truth | Operator re-types from Slack |
businessId | Search under pressure | Dump of executions nobody opens |
errorClass | Drives redrive policy | Every item looks the same |
partialState | Resume, do not restart | Duplicate writes |
owner | Names a human | Everyone assumes someone else |
idempotencyKey | Makes replay safe | Second charge / second contact |
status | Queue, not archive | Items sit open forever |
Views that get used:
- Open, sorted by age
- SEV1 / customer-facing only
- Needs schema fix (grouped by
errorMessage) - Auth / paused workflows
- Stale
in_progress(someone started and walked away)
Redact tokens, cookies, and any PII you do not need to replay. A DLQ that stores live API keys is an incident with a table UI.
How do you design replay so you do not double-write?
Replay is where teams re-introduce doubles. A “Replay” button that re-runs the entire production workflow is a footgun with a label.
Guardrails:
- Honor the idempotency key on every write node the replay can touch.
- Resume from the failed step when earlier steps already wrote.
- Mark the row
replayedwith timestamp and operator name before you fire, or in the same transaction if you can. - Never auto-replay schema or poison-payload failures until a human changes the payload or the contract.
- Cap automatic redrive for
transient_exhausted(three attempts, business hours, then human). - Do not redrive from Slack. The button lives on the row.
| Replay style | When it is legal | When it duplicates |
|---|---|---|
| Full graph from webhook | No side effects yet, or every write is idempotent | After any create/charge/send |
| Resume node / skip-create | partialState has the ids | You ignore those ids |
| Upsert by external id | CRM/email identity is stable | You key on a generated UUID each run |
| Compensating delete + restart | Policy allows it and the delete is safe | Finance, messaging, or anything irreversible |
| Bulk redrive after vendor outage | Class is transient_exhausted, concurrency capped | Schema items sneaked into the same view |
- Idempotency key present on the row
- Partial ids recorded if any write succeeded
- Replay path named on the workflow doc (
resumevsfull) - Operator name written on the row
- Auto-redrive disabled for
schemaandauth
If you cannot explain the replay path in one sentence, you are not ready to auto-retry.
What happens when a partial apply already wrote?
Example: step 1 creates a HubSpot contact, step 2 fails to enroll a sequence.
Blind retry creates a second contact unless step 1 upserts. The DLQ row must carry contactId and a resume instruction that skips create.
Correct paths:
- Upsert by email / external id on every create.
- On failure after create, write DLQ with
partialState.contactIdanderrorClass=partial. - Resume node: load
contactId, skip create, retry only the failed step. - Compensating delete only when policy allows and the delete cannot orphan a billed object.
| Step outcome | Next failure | Illegal retry | Legal park |
|---|---|---|---|
| Contact created | Sequence enroll fails | Restart from webhook | Resume enroll with contactId |
| Invoice drafted | Send email fails | Recreate the invoice | Send using invoiceId |
| Payment captured | Receipt email fails | Recapture | Send receipt only |
| Row upserted | Downstream webhook 500 | Re-upsert in a loop | Park; retry webhook with same key |
| Nothing written | First HTTP 503 | Unbounded retry | Bound retry, then transient_exhausted |
Document partial-state handling for every multi-write workflow. If you cannot explain it, turn the graph off until you can. Bravery is not a restore strategy.
How do you write a redrive policy before the first incident?
Write the policy before the first 429. On-call should not invent ethics at 11pm.
errorClass | Auto-redrive? | Human action | Then |
|---|---|---|---|
schema | Never | Fix contract or repair payload | Manual resume |
auth | Never | Pause workflow, rotate credential | Bulk replay with care |
partial | Never full-restart | Resume path only | Mark replayed |
transient_exhausted | Yes, up to N, business hours, concurrency cap | If N fails, human | Confirm vendor status page first |
unknown | Once, maybe | Classify it | Do not leave it unknown twice |
Vendor-outage playbook:
- Expect transient retries to burn their budget.
- Overflow to DLQ with
errorClass=transient_exhausted. - Post one status note in the alert channel (“vendor outage, redrive after 15:00”).
- Bulk redrive when the status page clears, with a concurrency limit.
- Confirm idempotency before the bulk button exists.
Without that playbook, every outage becomes twenty people pressing Replay differently.
SQS calls this a redrive policy. Service Bus calls it resubmit. Pub/Sub calls it a second subscription. Your Airtable view needs the same words written down.
How do you operate the queue so it is not a junk drawer?
A DLQ without a cadence is a folder named misc. Slack without a table is worse: the payload dies in scrollback.
Daily: triage new SEV1 / customer-facing items.
Weekly: close or discard stale SEV3; fix systemic schema issues.
Monthly: report top failing workflows to whoever owns roadmap time.
| Severity | Meaning | Eyes | Channel |
|---|---|---|---|
| SEV1 | Money movement or customer message may be wrong or missing | One business hour | Never-mute |
| SEV2 | CRM state wrong, internal impact | Same day | Morning triage |
| SEV3 | Enrichment / reporting gap | Next business day | Weekly board |
Pick numbers that match the business. Publish them. Hit them. If everything is SEV1, nothing is.
Ownership models:
| Model | Works when | Breaks when |
|---|---|---|
| Workflow-owner | SMB, <20 graphs | One person drowns |
| Domain on-call | Sales vs finance vs ops | Metadata has no owner |
| Central automation ops | Many graphs, shared platform | They triage but never assign out |
Put owner on the row from workflow metadata. Do not wait for humans to self-assign during an outage.
Healthy metrics:
- SEV1 open count near zero at start of day
- Median age under the SLA
- Recurring schema errors trending down after a contract fix
- Redrive success high for
transient_exhausted, near-zero auto-redrive forschema
A permanently non-empty DLQ is a product backlog. Schedule fix time. Do not celebrate depth.
Alert hygiene so the queue stays visible:
- Rate-limit Slack (burst of 200 failures → one summary + filtered view).
- Separate
#auto-criticalfrom#auto-noise. - Include
errorClassso people can ignore a known vendor outage. - Auto-close discarded items after the retention window, once finance has the export they need.
Alert fatigue fills DLQs as surely as missing alerts do.
What anti-patterns fill a DLQ forever?
These are the ones that show up after the first quiet month.
| Anti-pattern | What it looks like | Replace with |
|---|---|---|
| Retry storm, no ceiling | 429s, shared credential dead | Bound + DLQ |
| DLQ as logging only | Rows, no views, no owner | Cadence + SLA |
| Slack as the queue | “I think I saw that payload” | Table with rawInput |
| Huge raw payloads with secrets | Tokens in the cell | Redact; store a pointer |
| One shared DLQ, no owner | Everyone waits | owner from metadata |
| Silent Continue on Fail | Graph “succeeds,” work vanishes | Error output → DLQ write |
| Auto-redrive of poison | Same schema error, hourly | Human repair first |
| Full-graph replay after partial write | Duplicate contacts / charges | Resume path |
| Missing Error workflow attachment | New graph ships naked | Checklist on publish |
| Staging pointed at the P1 channel | Handler gets muted | Quiet handler, same fields |
Cross-workflow poison: the graph that threw is not always the graph that lied. When DLQ volume spikes on “missing email,” inspect yesterday’s enrichment writer. Keep a short dependency map of which workflows write fields others require.
Approvals are intentional waits. DLQs are broken waits. Keep separate tables or clearly separated statuses. Mixing “waiting on CFO” with “schema invalid” trains people to ignore the queue. A timed-out approval is escalation policy, not a DLQ write, unless the timeout should fail a downstream system.
How do you drill the failure path in staging?
If you have never failed it on purpose, you do not have a DLQ. You have a hope.
Once a quarter in staging:
- Send a payload missing a required field → expect
errorClass=schema, a row, and an alert. No vendor retries. - Force a 500 from a mock API → expect bounded retries, then
transient_exhausted. - Create a partial apply (mock CRM success, email fail) → expect
partialStateand a resume path, not a second CRM row. - Force a 401 → expect the workflow to pause, not to spray retries across the shared credential.
- Replay each case through the documented path. Confirm
replayedByand no duplicate side effects. - Confirm Execute Workflow in the editor does not fire the Error Trigger, then prove it on an activated Schedule or Webhook.
| Drill | Pass | Fail |
|---|---|---|
| Missing field | One DLQ row, zero extra API calls | Retry loop or silent Continue |
| Mock 500 | N retries, then park | Unbounded or no row |
| Partial apply | Resume uses the created id | Second create |
| 401 | Workflow paused | Shared 429 |
| Editor execute | Handler silent | False confidence |
| Activated path | Handler + row | “We thought it was attached” |
If the drill fails, production will fail louder. Attach the handler again after every duplicate or import — copied workflows often arrive with a blank Error workflow field.
The automation lane is where we install this as a standard, not a one-off Slack node.
FAQ
What is a dead letter queue for automation?
A durable place to store failed work with its input and error context so a human can fix and replay it. In n8n that is usually a table plus an error workflow and an alert — not a formal message broker. The broker docs (SQS, RabbitMQ, Service Bus, Pub/Sub) are the contract you copy: bound attempts, isolate the body, keep a replay path.
How should an n8n error workflow work?
On failure, capture execution id, failed node, error, and the original input (persisted earlier or fetched — the Error Trigger does not include the payload). Write a DLQ record, notify the owner with a deep link and a next action, and leave the item open until it is replayed or discarded. Do not pretend success, and do not let the handler call the same vendor again.
Should every failure go to the DLQ?
Transient failures should retry first with a bound and a wait longer than the vendor window. After that budget, or on poison, schema, or auth failures, yes. Read-only enrichment misses can log-and-continue if the business accepts the gap. Continue on Fail without a DLQ write is how work disappears.
How do I replay safely?
Replay through a path that checks the idempotency key and understands partial state. Prefer resuming after successful steps. Record who replayed and when. Do not auto-redrive schema or auth items, and do not re-run the whole production graph because Slack had a button.
Is a Slack message enough instead of a DLQ?
No. Slack is the alert. Without a stored payload, status, and owner, you cannot reconstruct or audit what failed, and you cannot replay without guessing. Use Slack to pull a human to the queue. The queue is the row.
What tool should store DLQ records?
Use whatever the team already queries: Postgres, Airtable, Notion. At Spurlock Studios we care that it is searchable by businessId, assignable, redactable, and replayable — not that it is fashionable. A broker DLQ is fine when the rail is already SQS or Service Bus; an n8n graph still needs the same fields.
CTA
If your error strategy is “it retries,” you do not have an error strategy.
Add a DLQ to the workflows that touch money or customers, keep the production handbook open while you wire it, and when you want this built as a standard, start at automation or book the $500 Automation Audit.
What questions does this article answer?
- What is a dead letter queue for automation?
- A durable place to store failed work with its input and error context so a human can fix and replay it. In n8n that is usually a table plus an error workflow and an alert — not a formal message broker. The broker docs (SQS, RabbitMQ, Service Bus, Pub/Sub) are the contract you copy: bound attempts, isolate the body, keep a replay path.
- How should an n8n error workflow work?
- On failure, capture execution id, failed node, error, and the original input (persisted earlier or fetched — the Error Trigger does not include the payload). Write a DLQ record, notify the owner with a deep link and a next action, and leave the item `open` until it is replayed or discarded. Do not pretend success, and do not let the handler call the same vendor again.
- Should every failure go to the DLQ?
- Transient failures should retry first with a bound and a wait longer than the vendor window. After that budget, or on poison, schema, or auth failures, yes. Read-only enrichment misses can log-and-continue if the business accepts the gap. Continue on Fail without a DLQ write is how work disappears.
- How do I replay safely?
- Replay through a path that checks the idempotency key and understands partial state. Prefer resuming after successful steps. Record who replayed and when. Do not auto-redrive schema or auth items, and do not re-run the whole production graph because Slack had a button.
- Is a Slack message enough instead of a DLQ?
- No. Slack is the alert. Without a stored payload, status, and owner, you cannot reconstruct or audit what failed, and you cannot replay without guessing. Use Slack to pull a human to the queue. The queue is the row.
- What tool should store DLQ records?
- Use whatever the team already queries: Postgres, Airtable, Notion. At Spurlock Studios we care that it is searchable by `businessId`, assignable, redactable, and replayable — not that it is fashionable. A broker DLQ is fine when the rail is already SQS or Service Bus; an n8n graph still needs the same fields.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation Why doesn’t worker concurrency cap my n8n sub-workflows
Worker concurrency does not cap n8n sub-workflows. Each Execute Workflow child is a new execution the production limit skips, usually on the parent worker.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.