n8n Error Workflows Operators Actually Read (Not Just Slack Noise)
Wire one n8n Error Trigger handler with a real alert contract: severity, owner, execution link, and mute rules — then attach it to every production flow.
William Spurlock Founder — Spurlock Studios Updated 18 MIN
An n8n error workflow people act on is a shared Error Trigger handler with a fixed alert contract — severity, owner, execution link, failed node, and next action — attached to every production workflow. A bare “post to Slack” node is how channels get muted.
Spurlock Studios treats the error handler as product surface, not an afterthought. Across 500+ automations, the graphs that stay trusted are the ones where a human can open the alert and know what to do in under a minute. Principles live in the Production n8n handbook; this spoke is the implementable contract.
The short answer
- One shared handler for production flows; set it under each workflow’s Settings → Error workflow.
- Alert contract beats clever copy — fields first, prose second.
- Manual Execute does not fire the Error Trigger; test with an activated automatic path or a mocked payload.
- Continue on Fail is for controlled branches, not a substitute for the error workflow.
- DLQ stores the work; the error workflow wakes a human — use both (dead letter queues).
What is an n8n Error Trigger workflow for?
n8n’s error-workflow docs are blunt: for each workflow you set an error workflow in Workflow Settings. It runs if that execution fails. The handler must start with an Error Trigger. You can point many production graphs at the same handler.
That is the whole job. The handler is not a debugger, not a retry engine, and not a place to “just send Slack.” It is the wake-up path when a linked workflow dies.
| Job | Owner | Not the handler’s job |
|---|---|---|
| Wake a human with a fixed field set | Error workflow | Dump raw JSON into #general |
| Deep-link the failed execution | Error workflow | Ask the operator to search Executions by memory |
| Confirm a DLQ write | Error workflow | Store the payload forever inside Slack history |
| Classify P1 vs P3 | Error workflow (table, not vibes) | Infer severity from the error string at 2am |
| Replay the item | DLQ + idempotency | Blind “retry execution” from the alert |
n8n also documents three facts that surprise people the first week:
- If a workflow uses an Error Trigger, you do not have to publish that handler for it to run when selected as an Error workflow.
- If a workflow contains an Error Trigger, it uses itself as its error workflow by default.
- You cannot test error workflows by clicking Execute Workflow in the editor. The Error Trigger only runs when an automatic workflow errors.
Error-workflow runs also do not count toward paid-plan execution quotas. n8n’s executions page lists them next to manual runs and empty polls as non-quota. That is not permission to spam the handler. It is permission to attach it without worrying the alert itself will burn your monthly cap.
What data does the Error Trigger receive?
The default payload in the Error Trigger reference looks like this (n8n’s own example shape):
| Field path | Present when | Use in the alert |
|---|---|---|
workflow.name / workflow.id | Almost always | Which flow broke |
execution.id | Execution was saved | Correlation id for DLQ and threads |
execution.url | Execution was saved | Deep link the operator clicks first |
execution.error.message | Mid-workflow failure | What failed, trimmed |
execution.lastNodeExecuted | Mid-workflow failure | Where to look first |
execution.retryOf | This run is a retry | Do not treat as a brand-new incident |
execution.mode | Mid-workflow failure | Context (automatic vs other) |
trigger.error.message | Trigger-node failure | Fallback copy when execution{} is thin |
trigger.error.node | Trigger-node failure | Fallback “failed node” |
Two caveats from the same page that matter in production:
execution.idandexecution.urlrequire the execution to be saved. They are missing if the trigger node of the main workflow failed, because that workflow never executed.execution.retryOfis present only when the execution is a retry of a failed execution.
If your Slack template assumes execution.url is always a string, the first webhook-activation failure will render a broken message. Broken templates train people to ignore the channel.
Normalize before you alert:
- Read
workflow.name(always try). - Prefer
execution.url; if missing, writeunavailable — trigger failureand setexecutionUrlAvailable: false. - Prefer
execution.lastNodeExecuted; fall back totrigger.error.node.name. - Prefer
execution.error.message; fall back totrigger.error.message. - If
execution.retryOfis set, tag the alertretry-ofso the first responder does not open a second incident.
Operators should never see an undefined field in Slack. Missing is a value. Write it.
What belongs in the alert contract?
Every P1/P2 alert must include the same lines, in the same order. Clever prose is optional. Structure is not.
severity: P1 | P2 | P3
workflowName
workflowId
executionUrl (or "unavailable — trigger failure")
failedNode
errorMessage (trimmed)
ownerPrimary
ownerBackup
customerOrRecordId (if known)
idempotencyKey (if any)
dlqStatus: written | skipped | n/a
nextAction: pause | replay | wait-retry | ignore-enrichment
occurredAt
If a field is unknown, write unknown — do not omit the line. Operators scan for missing structure faster than they read paragraphs.
| Field | Why it is there | Failure if omitted |
|---|---|---|
severity | Routes pager vs morning triage | Everything looks like P1, then nothing does |
executionUrl | One click into the failed run | Five minutes of searching Executions |
failedNode | Starts the debug at the right box | Scroll-and-guess |
ownerPrimary / ownerBackup | Names a human | “Engineering” owns it; nobody acts |
dlqStatus | Confirms the work was parked | Alert without a replay handle |
nextAction | Tells the first responder what to do | Thread of opinions, no pause |
nextAction is the field most teams skip and the one that makes the alert readable. Pause means stop the workflow before the next cron. Replay means take the DLQ row, not the Slack screenshot. Wait-retry means a bound backoff is already in flight. Ignore-enrichment means this is P3 and the pager was a mistake.
Do not put stack traces in the first message. Trim errorMessage to one line. Link the execution for the rest.
How do you map severity so a human can act?
Reuse the overnight posture from when automation fails overnight: severity is a route, not a feeling.
| Severity | Route | Example | First human action |
|---|---|---|---|
| P1 | Pager / SMS + never-mute channel | Payment, CRM overwrite, customer message | Pause the workflow, then open the execution |
| P2 | Morning triage channel | Lead sync lag, reporting job | Queue it for the next working block |
| P3 | Weekly board | Optional enrichment skip | Do not page; log and move on |
Map severity inside the error workflow with a simple table on workflow name, id, or tag. Do not make humans infer it from the error string at 2am.
workflowId → severity
wf_billing_invoice → P1
wf_crm_contact_upsert → P1
wf_customer_sms → P1
wf_lead_enrichment → P3
wf_weekly_report → P2
default → P2
Default to P2, not P1. An unknown workflow that pages phones will get the handler muted in a week. An unknown workflow that lands in morning triage gets a human who can raise it.
- Money / billing paths are P1
- CRM create/update paths that can overwrite truth are P1
- Customer messaging paths are P1
- Nightly reconciliation is P1 if a miss ships wrong numbers to a customer, else P2
- Enrichment, scoring, and “nice to have” fetches are P3
- Staging never pages a phone
If you cannot fill that list, you do not have a severity model. You have a Slack integration.
How do you wire one handler and attach it everywhere?
n8n’s setup steps are the same on the Error Trigger page and the workflow settings page: create the handler, save it, then on each production workflow open Options → Settings and pick it under Error workflow.
Procedure:
- Create workflow
Error Handler — Productionwith Error Trigger first. - Build: normalize payload → classify severity → write DLQ row → send alert → (optional) update the same thread on duplicates.
- Save. Confirm it appears in the Error workflow dropdown.
- For each production workflow: Options → Settings → Error workflow → Error Handler — Production → Save.
- Keep a checklist of attachments. New workflows do not inherit this by magic.
Attachment checklist:
- Money / billing paths
- CRM create/update paths
- Customer messaging paths
- Nightly reconciliation crons
- Webhook receivers that acknowledge early then process
- Any graph that can send email or SMS
- Any graph that deletes or merges records
Staging can share a quieter handler that never pages phones. Same field contract. Different channel. Do not “save time” by pointing staging at the P1 channel.
A workflow that contains an Error Trigger uses itself as its error workflow by default. That is fine for a dedicated handler. It is a foot-gun if you drop an Error Trigger onto a billing graph “to test” and forget it is now self-handling instead of calling the shared contract.
Re-check the dropdown after a duplicate or import. Copied workflows often arrive with a blank Error workflow field even when the source graph was attached.
Why didn’t the error workflow fire on a manual run?
Because n8n says it will not. The Error Trigger docs: you cannot test error workflows when running workflows manually. The node only runs when an automatic workflow errors.
Automatic, in n8n’s execution-mode language, means a published production run — Schedule, Webhook, or another trigger that fires without the editor’s Execute Workflow button. Manual editor runs are a different mode. They show the error in the canvas. They do not start the handler.
| How you failed it | Error Trigger fires? | What you actually tested |
|---|---|---|
| Editor Execute Workflow | No | Node settings and the red error on the canvas |
| Editor Execute step on one node | No | That node |
| Published Schedule Trigger + forced fail | Yes | The real attach path |
| Published Webhook production URL + forced fail | Yes | The real attach path |
| Mock Error Trigger JSON inside the handler | No (by design) | Alert formatting and DLQ write only |
Schedule Trigger has its own gotcha: n8n tells you to save and publish the workflow or the schedule does not run. A draft cron that you “run once” from the editor is still a manual execution.
Webhook has two URLs. The test URL is for Listen for Test Event. The production URL registers when you publish. If you POST to the test URL while staring at the editor, you are still in the manual/test path. Use the production URL on a published throwaway graph when you want the Error Trigger.
If you only clicked Execute in the editor, you have not tested the Error Trigger.
Continue on Fail vs Error Workflow — which when?
These are different tools. Mixing them is how errors go silent.
n8n’s node settings expose On Error as three choices: Stop Workflow, Continue, and Continue (using error output). Retry On Fail is a separate toggle: when a node fails, n8n reruns it until it succeeds. That last sentence is why Retry On Fail without a bound is a stampede, not a safety net.
| Mechanism | Use when | Avoid when |
|---|---|---|
| Error Workflow | The run should fail closed and a human/DLQ path must run | You want the item to continue downstream |
| On Error → Continue | A specific node may fail and you already have a branch | You enable it globally to “keep going” |
| On Error → Continue (using error output) | You will handle the error item on the error output | You ignore that output |
| Retry On Fail | Transient network / 429 with a budget you can name | Validation, auth, or poison payloads |
| Stop and Error | You want a controlled failure message into the Error Trigger | Debugging only in manual mode and expecting the handler to fire |
Continue on Fail without a branch that dead-letters or skips intentionally swallows API errors. That is how silent corruption starts. The CRM node fails, the next node writes yesterday’s item, and nobody gets a page because the execution did not fail.
Stop and Error is the honest opposite. n8n’s node page says it displays a custom error, fails the execution under your conditions, and sends that information to error workflows. Use Error Message when a one-line reason is enough. Use Error Object when you need structured fields the handler can read. Example: validator failed, you refuse to continue, you throw SCHEMA_INVALID instead of letting a later HTTP node throw a 400 with a vendor stack.
Still remember: Stop and Error during a manual Execute will not exercise the Error Trigger. Prove it on an activated automatic path.
Decision list:
- Will a failure here corrupt money, CRM, or a customer message? → Stop the workflow. Let the Error Trigger fire. Write DLQ.
- Is this optional enrichment? → Continue (error output) → log P3 → do not page.
- Is this a 429 or a 503 you have seen recover in under a minute? → Retry On Fail with a named max, then fail closed.
- Is this a schema / auth / “this payload is poison” case? → Stop and Error with a clear message. No retry.
How do you test without trusting a green checkbox?
Four rungs. Skip the last two and you shipped a formatter, not a handler.
- Mock path. Temporarily put a Set / Edit Fields node with sample Error Trigger JSON in front of your alert and DLQ nodes. Execute the handler workflow. This proves the template and the DLQ write. It does not prove attach.
- Activated failure. Publish a throwaway workflow that uses Schedule or Webhook. Point its Error workflow at your handler. Force a failure (bad URL, Stop and Error). Invoke it automatically — production webhook URL or a published cron — not Manual Execute.
- Trigger-failure case. Break a webhook or cron activation path once. Confirm the template survives missing
execution.urland still names the workflow. - Mute drill. Fire five identical errors. Confirm collapse / dedupe still leaves one actionable message with a count, not five threads.
| Rung | Proves | Does not prove |
|---|---|---|
| Mock JSON in the handler | Template + DLQ shape | Settings → Error workflow attach |
| Published Schedule / Webhook fail | Attach + automatic fire | Trigger-node missing-url shape |
| Trigger-node failure | Missing execution.url copy | Duplicate collapse |
| Five identical errors | Dedupe / thread update | Owner map accuracy |
n8n will let you load a failed execution back into the editor with Debug in editor. That is how you fix the production graph after the alert. It is not how you test the handler. Debug-in-editor is a manual run of the failed workflow, which again will not fire the Error Trigger.
How do you stop the channel from getting muted?
Alert quality dies when volume is undifferentiated. The Slack node can Send a message and Update a message. Use both. First failure sends. Duplicates update the same thread with a count.
Mute-prevention rules:
- No enrichment skips on P1 channels.
- Collapse duplicates: same
workflowId+failedNode+ normalized message within 30–60 minutes → update count, do not open a new thread. - Separate channels for P1 vs triage. One channel cannot be both a pager and a scrapbook.
- Never
@channelon P3. Almost never@channelon P2. - Include
nextActionso the first responder knows whether to pause or wait. - Do not paste the full stack in the first message. Link the execution.
- Do not @ people for P3. A weekly board exists for a reason.
| Anti-pattern | What operators do | Fix |
|---|---|---|
Raw JSON to #general | Mute the channel | Contract fields, dedicated channels |
| One channel for all severities | Mute after the first 429 weekend | P1 never-mute + P2 triage |
| New thread per retry | Ignore the fifth one | execution.retryOf + Update message |
@channel on enrichment skip | Disable mentions | P3 stays off the pager |
Alert with no nextAction | Debate in-thread | Pause / replay / wait-retry / ignore |
If the handler is noisier than the failures, operators will mute the handler — not fix the workflows. That is the whole product problem. You did not fail to “do Slack.” You trained the team that automation alerts are decoration.
Store a short-lived key for collapse: workflowId|failedNode|normalizedMessage. Look it up before Send. If it exists and is younger than your window, Update. If not, Send and write the key. Keep the window in the 30–60 minute range. Overnight storms should still page once, then count.
Where does a dead-letter queue sit next to the alert?
Alerts without storage create panic. Storage without alerts creates a quiet pile. Build the pair. Do not re-litigate DLQ theory here — that is the dead-letter queue post.
| Concern | Owner |
|---|---|
| Wake a human with context | Error workflow alert |
| Store payload + error for replay | DLQ table / queue |
| Prevent duplicate side effects on replay | Idempotency keys |
| Decide retry vs park | Failure classification |
| Confirm the park happened | dlqStatus on the alert |
Write the DLQ row before or as you send the alert, not after. If Slack is down, you still have the work. If the DLQ write fails, the alert must say dlqStatus: skipped so the human knows the screenshot is the only copy.
Minimum DLQ columns the handler should write when the original input is available:
-
workflowId/workflowName -
executionId(orunknownon trigger failure) -
failedNode -
errorMessage - original item / payload
-
idempotencyKeyif the main graph had one -
receivedAt -
status: parked
Replay from that row. Do not replay from a Slack message. Slack is the doorbell. The DLQ is the work order.
How do you keep the execution link from going missing?
execution.url is not a courtesy field. It is the difference between a one-click debug and a five-minute hunt. n8n only includes it when the execution was saved.
On each production workflow, open workflow settings and confirm:
| Setting | Production default I want | Why |
|---|---|---|
| Error workflow | Shared production handler | Attach is not inherited |
| Save failed production executions | On | execution.id / execution.url need a saved run |
| Save successful production executions | On until you have a reason | Replay and audit; prune later if storage hurts |
| Save execution progress | Off unless you need mid-node resume | Extra latency; not required for the alert link |
| Timeout Workflow | On, with a named budget | Hung HTTP should fail closed into the handler |
| Timezone | Set on Schedule graphs | Cron “didn’t fire” bugs are often timezone bugs |
Self-hosted instances can also set the same save policy globally. n8n’s execution environment variables include EXECUTIONS_DATA_SAVE_ON_ERROR (all / none, default all). If someone set that to none to “save disk,” your Error Trigger will still fire, but the alert will have no deep link and Debug in editor will have nothing to load.
Trigger-node failures are the remaining hole. The main workflow never executed, so there is no execution to save. Your template must say so. Pair that with a separate “did the cron fire?” check — a heartbeat row the Schedule graph writes on success — because a trigger that never ran will not page through this handler at all. That detection problem lives in the overnight post, not here.
Redaction is a second hole. If the instance redacts production execution data, the operator can still open execution.url and see status, timing, and node names, but not the payload. The DLQ is then the only place the item still exists. That is another reason dlqStatus belongs on the alert.
Failure mode: “we set up Slack” theater
What breaks: one Error Workflow posts raw JSON to #general with the Slack Send operation and no field contract. After a rate-limit weekend the channel is muted. A billing workflow fails on Monday. Nobody sees it until a customer asks.
What it costs: missed collections, manual invoice repair, and a team that no longer trusts automation alerts. The next failure gets a shrug. The handler is still “green” in the editor because nobody ever tested it automatically.
What you do instead:
- Ship the contract above — same fields, same order,
unknownwhen empty. - Attach the shared handler to every production flow. Checklist, not memory.
- Prove an automatic failure once (published Schedule or production Webhook).
- Prove the missing-
execution.urlrender once. - Collapse duplicates. Split P1 from triage.
- Write the DLQ row and put
dlqStatuson the message. - Protect the P1 channel like a pager. Enrichment never lands there.
This is the failure I see after the first 500+ automations more than any missing node: the team built a notification and called it operations. Notifications that cannot be acted on are noise. Noise gets muted. Muted P1 is just a delayed outage.
What does the production attach runbook look like?
Copy this. Do not ship the handler until step 6 is checked.
1. Handler workflow saved and named (Error Trigger first)
2. Alert template includes all contract fields
3. DLQ write step verified with mock payload
4. Severity routing table filled for this app domain
5. Each production workflow Settings → Error workflow set
6. Automatic failure test passed (not manual Execute)
7. Trigger-node missing-url case rendered safely
8. Duplicate collapse verified (Send then Update)
9. Owners + backups named in the alert body
10. Save failed production executions is on
11. nextAction is one of: pause | replay | wait-retry | ignore-enrichment
12. P1 channel is not the enrichment channel
13. Link to pause steps in the team runbook
14. Staging points at the quiet handler
First-five-minutes card for the human who got paged:
| Minute | Action |
|---|---|
| 0 | Read severity, nextAction, ownerPrimary |
| 1 | If nextAction is pause, unpublish / deactivate the workflow |
| 2 | Open executionUrl or, if unavailable, treat as trigger failure |
| 3 | Confirm dlqStatus. If skipped, you are working from the alert only |
| 4 | Decide: replay from DLQ, wait a bound retry, or leave it parked |
| 5 | Write the outcome on the thread. Do not start a second incident |
Skip step 6 on the runbook and you are shipping hope. Green in the editor is not an error workflow. An error workflow is a human who can act.
Owner fields belong in the alert body. Do not send people to a wiki. Source ownerPrimary and ownerBackup from a static map in the handler (workflow id → people), a small lookup table, or a naming convention you read carefully. Static maps drift when roles change; revisit them when someone leaves. A wrong owner still gets corrected. An empty owner gets ignored.
Pause, in this runbook, means unpublish or deactivate the failing workflow so the next cron or webhook cannot double-apply. It does not mean “leave a Slack reaction and hope.” Replay means the DLQ row, with the idempotency key checked, not a blind retry of the whole execution from the editor.
FAQ
Why didn’t my error workflow fire on a manual run?
Because n8n does not run the Error Trigger on manual executions. The docs are explicit: it only runs when an automatic workflow errors. Activate a test workflow and fail it via a published webhook or schedule, or mock the payload inside the handler to test formatting. An editor Execute that shows a red node has not tested attach.
Should every workflow share one handler?
Share one production handler for a consistent alert shape and DLQ write. Use a separate quiet handler for staging that never pages phones. Special-case a second production handler only when a domain truly needs a different pager route — and still keep the same field contract. New workflows do not inherit the setting; attach it on each graph.
How do I log failures for replay?
In the error workflow, write the original input (when available), error, execution id, and workflow id to your DLQ store before or as you alert. Replay from that record with idempotency checks, not from a Slack screenshot. The field list and replay rules live in the dead-letter queue guide.
Continue on Fail vs Error Workflow — which when?
Use the Error Workflow when the execution should fail closed and notify. Use Continue on Fail only on a specific node with an explicit branch that skips, repairs, or dead-letters. Do not enable Continue (or Continue using error output) to hide errors. Retry On Fail is for transient network and 429s with a named budget, not for validation or auth.
How do I avoid paging for enrichment skips?
Classify enrichment workflows as P3, or handle skips inside the main flow with Continue on Fail plus a non-pager log. Only money, CRM integrity, and customer-contact failures belong on the phone. If enrichment shares a channel with P1, operators will mute the channel and you will miss the invoice failure.
Where does DLQ fit next to alerts?
DLQ holds failed work for repair and replay. The error workflow notifies humans and should confirm the DLQ write with dlqStatus. Alert without storage is a screenshot culture. Storage without alert is a forgotten queue. Build both; do not pick one and call it production.
CTA
If your Error Trigger only dumps JSON into Slack, you built a mute button. Ship the contract, attach it everywhere, and prove an automatic failure once.
Review the spine in the handbook, then use automation or book the $500 Automation Audit.
What questions does this article answer?
- Why didn't my error workflow fire on a manual run?
- Because n8n does not run the Error Trigger on manual executions. The docs are explicit: it only runs when an automatic workflow errors. Activate a test workflow and fail it via a published webhook or schedule, or mock the payload inside the handler to test formatting. An editor Execute that shows a red node has not tested attach.
- Should every workflow share one handler?
- Share one production handler for a consistent alert shape and DLQ write. Use a separate quiet handler for staging that never pages phones. Special-case a second production handler only when a domain truly needs a different pager route — and still keep the same field contract. New workflows do not inherit the setting; attach it on each graph.
- How do I log failures for replay?
- In the error workflow, write the original input (when available), error, execution id, and workflow id to your DLQ store before or as you alert. Replay from that record with idempotency checks, not from a Slack screenshot. The field list and replay rules live in the [dead-letter queue guide](/blog/dead-letter-queues-for-automations).
- Continue on Fail vs Error Workflow — which when?
- Use the Error Workflow when the execution should fail closed and notify. Use Continue on Fail only on a specific node with an explicit branch that skips, repairs, or dead-letters. Do not enable Continue (or Continue using error output) to hide errors. Retry On Fail is for transient network and 429s with a named budget, not for validation or auth.
- How do I avoid paging for enrichment skips?
- Classify enrichment workflows as P3, or handle skips inside the main flow with Continue on Fail plus a non-pager log. Only money, CRM integrity, and customer-contact failures belong on the phone. If enrichment shares a channel with P1, operators will mute the channel and you will miss the invoice failure.
- Where does DLQ fit next to alerts?
- DLQ holds failed work for repair and replay. The error workflow notifies humans and should confirm the DLQ write with `dlqStatus`. Alert without storage is a screenshot culture. Storage without alert is a forgotten queue. Build both; do not pick one and call it production.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation Why doesn’t worker concurrency cap my n8n sub-workflows
Worker concurrency does not cap n8n sub-workflows. Each Execute Workflow child is a new execution the production limit skips, usually on the parent worker.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.