Spurlock Studios
Contact
Share LinkedIn X
Amber node beads on a dark rail. Thesis: CONTINUE FAIL HIDE REAL API.

Continue on Fail hides real API errors in n8n when a node throws and the execution still succeeds. From n8n’s point of view the run handled it, so the Error Workflow never starts. Downstream nodes keep going — and on On Error → Continue, they often keep going with the last valid data, which is how a 401, a 429, or a vendor 500 becomes a quiet wrong write.

This is not a settings nit. It is the difference between a red execution you can replay and a green execution that already emailed a customer, merged a CRM record, or skipped the only enrichment that made the row true. The production spine lives in the Production n8n handbook. This spoke is the Continue trap: when the toggle is honest, when it lies, and how to prove it before staging promotion.

The short answer

  • Continue hides errors when the execution stays green. Error Trigger requires a failed run. A continued node is not a failed run.
  • On Error → Continue uses last valid data. n8n’s node settings say so. The next write can apply the previous item.
  • On Error → Continue (using error output) is the only Continue that hands you the error. If that pin is unconnected, failed items vanish and the run still succeeds.
  • Retry On Fail is a bound for transient HTTP. After Max Tries, fail closed. Do not Continue the leftovers on a money path.
  • HTTP Never Error is a second mute. 4xx/5xx return as success, so On Error never fires. Branch on statusCode or do not turn it on.

What does Continue on Fail actually do in n8n?

It changes whether a node failure is allowed to fail the execution. The editor still calls the family “Continue on Fail.” Current Settings copy is On Error.

n8n documents three choices on the node Settings tab:

On Error valueWhat n8n doesExecution resultError Workflow
Stop WorkflowHalts. No further nodes.FailedFires (automatic runs only)
ContinueNext node runs using last valid dataSuccessDoes not fire
Continue (using error output)Error info goes to the error output pinSuccess unless you later StopDoes not fire unless you fail it yourself

Two other Settings sit next to that dropdown and get mixed in:

SettingOfficial behaviorTrap
Retry On FailNode reruns on failureUnbounded “until it succeeds” on a write is a stampede
Always Output DataEmpty item if the node returned nothingEmpty item looks like a successful skip; IF nodes can loop

Continue is not “log the error and keep the payload.” Continue (the first option) is “pretend this node produced usable data.” That is the hide.

Decision list:

  1. Will a failure here corrupt money, CRM, or a customer message? → Stop Workflow. Let the Error Trigger fire.
  2. Is this optional enrichment you can name as skippable? → Continue (using error output) → log → do not page.
  3. Is this a 429 or a 503 you have seen recover? → Retry On Fail with Max Tries and Wait Between Tries, then Stop.
  4. Is this schema, auth, or poison? → Stop And Error. No retry. No Continue.

If you cannot say which row you are on, you are on Continue-as-mute. Turn it off.

Exported workflow JSON is the audit trail when the UI label moved. Keys have changed across n8n versions, so grep more than one name. As of the current editor, treat these as the family — confirm on a node you just toggled, because n8n can rename fields without a blog post:

What you set in SettingsWhat to grep in JSON (hedge: confirm on your version)Meaning
On Error → Stop WorkflowonError stop / missing Continue flagsExecution may fail
On Error → ContinuecontinueRegularOutput or legacy continueOnFail: trueLast valid data; run succeeds
On Error → Continue (using error output)continueErrorOutputError pin; run succeeds unless you Stop later
Retry On FailretryOnFail, maxTries, waitBetweenTriesBudget lives here — empty max is the stampede
Always Output DataalwaysOutputData: trueBlank item, not an error

Do not argue from memory of the old checkbox. Open one node, flip On Error, re-export, and match the key your instance actually writes. Then grep the rest of the instance for that key.

A batch of items does not make Continue safer. One node that processes 50 leads with Continue still has one “last valid” at a time. Isolate with Loop Over Items (or equivalent) and Stop on the write inside the loop, or you will still donate lead 46’s ids to lead 47’s failure.

When does the Error Workflow stay silent?

Always, if the execution did not fail. n8n’s error-workflow docs are one sentence long: for each workflow you set an error workflow in Workflow Settings; it runs if that execution fails. Continue means it did not fail.

The handler must start with an Error Trigger. Attach it under Options → Settings → Error workflow. New graphs do not inherit it.

What you didExecution statusError Trigger
Node errors, On Error = Stop Workflow, automatic runErrorYes
Same node, On Error = ContinueSuccessNo
Same node, Continue (using error output), error pin ignoredSuccessNo
HTTP Never Error on a 500SuccessNo
Editor Execute Workflow even with Stop WorkflowError on canvasNo — docs: Error Trigger only on automatic failures
Trigger node itself failed to activateOften no saved executionThin trigger.error payload, or nothing you expected

Silence is not “the API was fine.” Silence is “n8n did not classify this run as a failure.”

If the only proof you have is a quiet Slack channel, you have not measured errors. You have measured whether executions stayed green. Those are different numbers. The overnight version of this gap — trigger never fired, nobody woke — is When Automation Fails at 2am.

Checklist before you trust “no alerts”:

  • Every production workflow has Settings → Error workflow set to the shared handler
  • Write nodes use Stop Workflow, not Continue
  • You proved attach on a published Schedule or production Webhook URL, not Manual Execute
  • HTTP nodes do not have Never Error unless a later node branches on status
  • Continue exists only on named enrichment nodes with a connected error pin

If any box is unchecked, the quiet channel is the bug.

Error-workflow executions do not count toward paid-plan execution quotas on n8n’s executions page. That is not a reason to Continue “to save runs.” It is a reason to attach the handler. You are not being charged extra for honesty.

Also remember: a workflow that contains an Error Trigger uses itself as its error workflow by default. Do not drop a test Error Trigger onto a billing graph “to see the payload.” You will re-enter the handler on the handler’s own failures in ways you did not plan, and n8n will not chain an error-workflow-for-the-error-workflow. Keep the handler boring: normalize, DLQ write, alert. If the handler’s Slack node 401s, you read process logs — not a second pager graph.

Handler designAllowedNot allowed
Shared Error Trigger + DLQ + alertYesPer-graph “also Slack” with no fields
Staging handler that never pages phonesYesStaging using the production P1 channel
Continue on the handler’s HTTPNoSwallowing the park path
Manual Execute to “test attach”NoClaiming the handler is broken

The Error Trigger also does not include the input that caused the failure. If Continue already succeeded the run, you do not even get the trigger. Persist payloads before risky writes, or you will debug from a green bar and a vendor screenshot.

Continue vs Continue (using error output) — which one lies?

Continue lies by default. Continue (using error output) only lies if you ignore the pin.

n8n’s wording is the whole case. Continue: proceeds to the next node despite the error, using the last valid data. Continue (using error output): continues, passing error information to the next node for potential handling.

ContinueContinue (using error output)
What the next success-path node receivesLast valid itemSuccessful items only (main output)
Where the failure livesNowhere structuredError output pin
Typical downstream bugWrite yesterday’s row / previous customerDrop the failed item with no record
Honest useAlmost never on HTTP that feeds a writeEnrichment skip, fallback vendor, DLQ branch
Dishonest use“Keep the graph green”Leave the red pin unconnected

The last-valid-data path is the one operators miss in reviews. Item 12 fails HubSpot. Item 11 is still sitting there as last valid. The next node creates a note, sends a mail, or patches a deal using item 11’s ids. The execution is green. The API error existed. Nobody got paged.

The unconnected-pin path is quieter. Failed items go to an output you never wired. They do not continue. They also do not fail the run. You “handled” them by deleting them.

Procedure to inspect a node:

  1. Open the node → Settings.
  2. Read On Error. If it says Continue, treat it as last-valid-data until proven otherwise.
  3. If it says Continue (using error output), follow the second output. If it goes nowhere, that is a drop.
  4. Open a recent success execution. Click the node. If you see an error object on a green run, Continue already swallowed it.
  5. Compare execution.status to the node’s output. Green plus $json.error is the hide.

Do not “fix” Continue by adding Always Output Data. Empty items are not errors. They are blanks that look like skips.

IF expressions after a Continue node have to name the hide, or the IF is theater:

What you are detectingTypical check (inspect a live item; field names move)If true
Continue left an error object$json.error exists, or $json.error.message is non-emptyLog / Stop And Error — do not write
HTTP Never Error left a bad status$json.statusCode not in 200–299 (only if you included headers/status)Same
In-band vendor error on HTTP 200Body has error, errors[], or ok: false you have seen from that APITreat as failure; do not trust 200
Always Output Data blankItem has no keys you requireSkip or Stop — do not patch CRM with {}
Last-valid-data suspectBusiness id ≠ current loop item’s idHalt the batch; that is contagion

If the IF is $json.error !== undefined but you used Continue (last valid data), there may be no $json.error on the next node. The next node received the previous success. The IF never fires. That is why last-valid Continue cannot be “fixed” with an IF on the success path. You need Stop, or the error output pin.

Retry On Fail is a budget, not a mute

Retry and Continue solve different failures. Mixing them is how a 401 gets hammered three times and then walked downstream as last valid data.

n8n’s public Settings blurb for Retry On Fail is blunt: when an execution fails, the node reruns until it succeeds. That sentence is why you never enable it on “create invoice” without a cap. The HTTP Request common issues page is more useful in the editor: enable Retry on Fail, set Max Tries, set Wait Between Tries (ms).

FailureRetry On FailThen
429 rate limitYes, Wait Between Tries above the vendor windowStop if still 429; do not Continue
502 / 503 / timeoutYes, small Max TriesStop; page if a write node
401 / 403 authNoStop; credential incident, not a blip
400 / schemaNoStop And Error; payload is poison
404 on optional lookupNoContinue (error output) → log skip
409 conflict on createNo (usually)Idempotency check, not a blind retry

Retries happen before On Error applies. If retries exhaust and the node still throws:

  1. Stop Workflow → execution fails → Error Trigger can fire.
  2. Continue → execution succeeds → Error Trigger stays dark, last valid data continues.
  3. Continue (using error output) → execution succeeds unless you Stop later → you must handle the error pin.

Retry without a max is a stampede. Retry then Continue on a write is a stampede that reports success.

Named budget you can defend:

  • Max Tries: a number you would say out loud to the person who owns the vendor bill (often 3–5 for reads; lower for writes).
  • Wait Between Tries (ms): longer than the documented rate-limit window, not “1000 because the docs used 1000 as an example.”
  • After the budget: Stop on writes. Branch on enrichment. Never “keep going” into money.

If you cannot name Max Tries, you do not have retries. You have hope plus latency.

Rate limits have a documented n8n path that is still not Continue. On 429, the HTTP node errors with The service is receiving too many requests from you — see handling API rate limits and the HTTP common issues page. Their fix is pause, then retry: Batching (Items per Batch + Batch Interval), or Retry On Fail with Wait Between Tries (ms) above the vendor window.

Procedure for a 429-prone read:

  1. Leave Never Error off so 429 is a real node error.
  2. Settings → Retry On Fail on. Set Max Tries. Set Wait Between Tries longer than the vendor’s documented window (their 1000 ms example is “one request per second,” not a universal default).
  3. On Error = Stop Workflow after that budget.
  4. If the vendor sends Retry-After in seconds longer than the Wait Between Tries field is comfortable with, do not pretend the field is a 30-minute sleeper — use a Wait node on a branch, or park to DLQ.
  5. Do not Continue the 429 into the CRM write.

Retry is how you survive a blip. Continue is how you survive a demo. Only one of those belongs on a write node.

HTTP Never Error is a second silent green

Continue hides node exceptions. HTTP Request Never Error hides status codes so they never become exceptions.

n8n: by default the node returns success only on 2xx. Turn Never Error on and it returns success regardless of the code. The run is green. On Error never runs. Retry On Fail may never see a failure. Your IF node has to read the body like an adult.

MuteWhere it livesWhat stays greenWhat you must do instead
On Error → ContinueNode SettingsExecutionStop, or error-output branch
On Error → Continue (using error output), pin unusedNode Settings + canvasExecutionWire the pin to log / DLQ / Stop And Error
Response → Never ErrorHTTP Request optionsThe HTTP node itselfInclude headers/status; branch on statusCode
Always Output DataNode SettingsA blank itemDo not treat blank as “vendor said OK”

Never Error is legal for one job: you want the 404/409 in the graph so you can IF on it. Example: lookup customer, 404 means create, 200 means patch. That is a branch. It is not “don’t bother me with 500s.”

Procedure if Never Error is on:

  1. Add Option → Response → Include Response Headers and Status.
  2. IF / Switch on statusCode (and do not assume the field name until you inspect one live item).
  3. 2xx → happy path.
  4. 404 / expected 409 → the skip or create branch you designed.
  5. 401 / 403 / 429 / 5xx → Stop And Error (or Retry only for 429/5xx before this node, with Never Error off).

If Never Error is on and there is no status IF, you built a JSON pipe that treats { "error": "insufficient_quota" } as a successful customer record.

In-band failures still happen on honest 200s. Never Error does not create those; it just makes the 4xx cousins look the same. Branch both:

HTTP layerExampleContinue / Never Error help?Correct move
Transport / DNS / timeoutECONNREFUSED, abortNo — this is a node errorRetry bound, then Stop
Status 401 / 403Expired tokenNoStop; credential incident
Status 429Rate limitNoRetry + Wait; then Stop
Status 5xxVendor downNoRetry small; then Stop
Status 404 you expectedMissing remote recordYes, if Never Error or error-output, then IFCreate or skip on purpose
Status 200 + { "ok": false }GraphQL/app error in bodyNoSchema/IF on body; Stop And Error
Status 200 + empty required fieldsPartial payloadNoValidator → Stop And Error

Timeouts are their own mute if you never set one. HTTP Request Timeout (ms) aborts waiting for response headers. An unset long hang is not Continue, but operators treat “still running” like health. Set a timeout on writes. Then let Stop Workflow fire when it aborts.

Failure mode: last-valid data writes the wrong record

What breaks: a production HTTP or app node is set to Continue so “one bad lead doesn’t stop the batch.” Lead 47’s API call 401s. Lead 46 is last valid. The next node — Slack, Gmail, HubSpot note, invoice draft — runs with lead 46’s ids. The execution is success. The Error Workflow is quiet. You find it when a human notices a duplicate note or a missing lead, not when n8n tells you.

Cost: cleanup on the system of record, plus a week of distrust in every green run. Across 500+ live automations the expensive part is not the red execution. Red is cheap. Green and wrong is the invoice you cannot unsend.

A second version: Continue (using error output) with no pin. Lead 47 vanishes. CRM stays on yesterday. Dashboard counts “success.” Overnight, the batch looks healthy while the queue of real failures is empty because nothing was stored. That is how automation fails overnight without a pager: the trigger fired, the graph finished, the work did not.

A third version: Retry On Fail on a non-idempotent POST, then Continue. Three creates, then a “success” that still used last valid data. Now you have duplicates and a hide.

Do this instead:

  1. Loop Over Items (or equivalent) so one lead cannot donate its data to the next.
  2. Stop Workflow on the write node. Pair with an Error Trigger handler operators can read.
  3. Persist the failing item to a replay table before or as you alert — Error Trigger does not include the original input.
  4. Replay from that row with an idempotency key, not from “retry execution” in the editor.

If the graph cannot survive one 401 without writing the previous person, Continue is not resilience. It is contagion.

Reconstruction when you already shipped the mute — do this without inventing a “typical incident count”:

  • List success executions in the window since Continue was enabled on the write path
  • For each, compare write-node input ids to the trigger/item ids for that run
  • Mismatches are last-valid-data candidates; export them before you “retry all”
  • Search the CRM/mail tool for duplicate notes, duplicate sends, or skipped ids from the same window
  • Turn Continue off on writes before you replay anything, or you will duplicate while investigating
  • Replay only from stored payloads + idempotency keys, never from “the green execution looked fine”

If you never stored inputs, reconstruction is vendor-side forensics. That gap is why the Error Trigger is not a DLQ. The hide deleted the work order.

When is Continue actually honest?

When the failed node is skippable by name, the error is visible on a branch, and a human-readable skip is stored somewhere that is not Slack scrollback.

Honest Continue is almost always Continue (using error output) plus a log. It is almost never the first Continue option.

Node jobContinue OK?Required branch
Optional firmographic lookupYesError pin → write enrichmentStatus: skipped → proceed
Optional Slack notify of an internal eventYesError pin → log; do not page
Fallback vendor after primary 404YesError pin → second HTTP; if that fails, Stop
CRM create / updateNoStop Workflow
Email to a customerNoStop Workflow
Stripe / invoice / payoutNoStop Workflow
Credential refresh / OAuth callNoStop; Continue swallows the only expiry signal
Schema validatorNoStop And Error

Checklist for an honest enrichment Continue:

  • The node is not on a money, CRM-integrity, or customer-contact path
  • On Error is Continue (using error output), not Continue
  • Error pin writes a skip record (sheet, table, or execution custom data) with workflow, node, item id, and trimmed error
  • Success path does not read last week’s enrichment as if it just succeeded
  • Skip is P3 / non-pager. It does not share the P1 channel
  • A weekly pass exists so “temporary skip” cannot last a quarter

If the skip record does not exist, you did not continue. You discarded.

Enrichment Continue is still a product decision. “We don’t need that field” is allowed. “The API failed and the graph stayed green so we assumed we don’t need that field” is how stale firmographics become sales truth.

How do you find swallowed errors in a green execution?

You look at node output on success runs. Failed executions are the ones you already know about.

Procedure:

  1. Filter Executions to Success for the workflow, last 7 days.
  2. Open several, not one. One lucky run proves nothing.
  3. Click every HTTP and app node that has Continue, Never Error, or Always Output Data.
  4. Hunt for $json.error, empty items, vendor error bodies, or statusCode outside 2xx.
  5. Diff that against the next write node’s input. If the ids do not match the item you think you are on, last-valid-data already happened.
  6. Count: green executions with an error-shaped payload. That count is your hidden-error metric. Do not invent an industry rate. Use your instance.
SignalMeansAction
Success execution, node panel shows error objectContinue swallowed itFlip to Stop or wire error output
Success, HTTP body is vendor error JSON, status 200API lied in-band; Never Error may also be onBranch on body + status; fail closed on writes
Success, node output empty, Always Output Data onBlank skipDecide skip vs Stop; do not write blanks
Success volume high, Error Trigger volume zero, writes existClassic muteStaging fail test; then fix On Error
Error Trigger volume high, all enrichmentWrong Continue polarityMove enrichment to error-output + P3 log

Do not use “error rate” from a dashboard that only counts failed executions. That dashboard cannot see Continue.

A cheap instance scan:

  • Export or click through production workflows
  • List nodes where On Error ≠ Stop Workflow
  • List HTTP nodes with Never Error
  • List write nodes downstream of those
  • Any write downstream of Continue is a defect until a human signs the skip

You will find more mutes in an afternoon than in a month of “we’ll watch Slack.”

How do you wire a continue branch that does not lie?

You treat the error output like a second workflow that must finish in a known state: skip, fallback, or fail closed.

Numbered graph (enrichment example):

  1. HTTP Request (lookup). On Error = Continue (using error output). Never Error off. Retry On Fail only if this vendor 429s in real life, with Max Tries named.
  2. Main output → IF: body has the fields you need → merge onto the item → continue.
  3. Error output → Set/Edit Fields: enrichmentStatus=skipped, enrichmentError={{ trimmed message }}, keep the business ids from the input of the HTTP node, not from “last valid.”
  4. Write that skip to your log table (same store you would use as a mini-DLQ).
  5. Merge skipped items back only if the rest of the graph can run without the enrichment. If it cannot, do not merge. Stop And Error.
  6. Downstream writes read enrichmentStatus. They never assume the lookup ran.

Numbered graph (write example — do not Continue):

  1. Validate schema. On failure: Stop And Error with a stable code (SCHEMA_INVALID).
  2. HTTP / app write. On Error = Stop Workflow. Retry On Fail only for 429/503, Max Tries low, Wait Between Tries above the vendor window.
  3. Settings → Error workflow = shared handler.
  4. Handler writes DLQ + alert with execution link, failed node, owner, nextAction.
  5. Replay from DLQ with idempotency, not from a green sibling item.
PatternNodesAllowed Continue
Enrichment skipHTTP → error pin → log → optional mergeError output only
Fallback vendorHTTP A error pin → HTTP B; B stops on failOn A only
Fail closedWrite → Stop → Error Trigger → DLQNone
Status branchHTTP Never Error plus status IF → Stop on 5xxNever Error only with the IF

If step 3 (skip write) fails, that is not “best effort.” A skip you cannot see later is Continue-as-mute with extra steps. Fail the skip-log loudly or you are back in the trap.

Keep business ids from $input / the node before the HTTP call, not from the HTTP output, on the error pin. The error output is the failure. It is not a clean copy of the lead. If you map email from the failed HTTP body you will write blank or vendor HTML into the skip log and you will not be able to replay.

Stop And Error belongs at the end of a branch that decided the item is poison:

  1. IF: schema invalid or 401/403 on a write path.
  2. Stop And Error → Error Message for a one-line reason (SCHEMA_INVALID, AUTH_EXPIRED).
  3. Error Object when the handler needs fields (record id, vendor, statusCode).
  4. That failure is a real execution failure. Error Trigger can fire. Continue is not in this picture.

Do not put Stop And Error and Continue on the same node. The Settings On Error of Stop And Error should stay Stop Workflow. You are throwing on purpose.

How do you prove this in staging before production?

A green editor run does not prove Error Trigger attach, and it does not prove Continue is safe. Prove both. Details on the promotion gate: staging n8n before production.

Four rungs:

  1. Force a write-node 401 on a throwaway copy (bad token, or Stop And Error). On Error = Stop Workflow. Publish. Hit the production webhook URL or a published Schedule. Confirm Error Trigger fires and the alert has execution.url (or an explicit missing-link line).
  2. Same graph, On Error = Continue. Repeat the 401. Confirm the Error Trigger does not fire and the next node ran. That is the hide, on purpose, once, in staging.
  3. Continue (using error output) with the pin wired to a skip log. Confirm the skip row exists and the write node did not run with last-valid ids.
  4. HTTP Never Error on a 500 fixture. Confirm you need a status IF or the graph “succeeds” with a vendor error body.
RungPassFail
Stop + automatic failHandler alert + DLQ rowManual Execute only; “we’ll attach later”
Continue 401Next node ran; no handlerYou still think Continue pages
Error-output skipSkip row; write skipped or markedUnconnected pin; missing item
Never Error 500Status IF Stopped the writeBody landed in CRM

Do not skip rung 2. Teams argue about Continue in the abstract. One staging 401 ends the argument.

Also prove the negative: enrichment Continue must not land on the P1 channel. If the skip log pages a phone, operators will mute the handler and you will miss the invoice failure next week.

Webhook URL mix-ups fake a pass. n8n’s Webhook node has a test URL (Listen for Test Event) and a production URL that registers when you publish. POSTing the test URL while the editor is open is still the manual/test path. The Error Trigger docs: you cannot test error workflows by running them manually. Use the production URL on a published throwaway graph.

Schedule Trigger has the same class of lie: save and publish or the cron does not run. A draft you “run once” from the editor is still manual. Rung 1 is not optional.

How you invoked itTests On Error / Continue behavior on canvas?Tests Error Trigger attach?
Execute WorkflowYes — you will see red vs green nodesNo
Execute stepThat node onlyNo
Test webhook URL + ListenCanvasNo
Production webhook URL, publishedYesYes
Published Schedule, wait for tickYesYes
Debug in editor from a failed runThe failed graph, as a manual runNo

If staging never used the production URL, you staged a screenshot.

Audit an existing instance for Continue traps

Do this on any n8n box that already “runs production.” Do not wait for a customer to find last-valid data.

  1. Inventory activated workflows. Note which have Settings → Error workflow set. Unset = silent on Stop, too.
  2. Open each activated graph. For every HTTP Request and app node, record On Error, Retry On Fail (Max Tries / Wait), Never Error, Always Output Data.
  3. Mark write nodes (CRM, mail, invoices, deletes, Slack to customers).
  4. Draw one line: Continue or Never Error upstream of a write. Each line is a defect.
  5. Sample 10 success executions per money workflow. Search node output for error objects and non-2xx.
  6. Fix writes to Stop Workflow first. Leave enrichment Continue only where the skip log already exists or you add it in the same change.
  7. Prove attach with one automatic fail per handler (shared handler means one proof, plus a spot check that new workflows actually select it).
  8. Write the rule in the runbook: Continue is a named skip. It is not the default.
FindingSeverityFix this week
Write node on ContinueP1Stop Workflow; Error workflow attached
Never Error into a write with no status IFP1Never Error off, or IF then Stop
Retry On Fail, no Max Tries, on POSTP1Cap tries or disable
Enrichment Continue, no skip logP2Error output → log
Error workflow unsetP1Attach shared handler
Handler never tested automaticallyP2Staging rung 1

Hire vs DIY lives here. If you cannot finish steps 1–4 in a sitting, you do not have an ops graph. You have a canvas. The $500 Automation Audit is the paid version of this list when the writes can hurt.

Do not turn Continue off globally as a “cleanup.” Enrichment graphs will start failing closed and paging for skips. Classify first. Then change On Error one node class at a time.

Export-and-grep procedure when clicking through the UI will miss a node:

  1. Download workflow JSON for every activated production graph (UI export or source control, without secrets).
  2. Search for Continue-family keys you confirmed on a toggle test: continueOnFail, continueRegularOutput, continueErrorOutput, retryOnFail, alwaysOutputData, and HTTP options that look like Never Error (neverError / response never-error flags — confirm the literal key on one node you just enabled).
  3. Search for write-ish node types you actually use (n8n-nodes-base.httpRequest feeding CRM, Gmail, Slack, Stripe, HubSpot, etc.).
  4. Any Continue-family hit on or immediately upstream of a write type is a P1 until a human signs a skip.
  5. File the grep output next to the Settings screenshot. The JSON is what will drift back after someone “keeps the demo green” next sprint.

Hedge: n8n’s export shape has moved. Do not treat a missing key as proof of Stop Workflow until you have the paired sample from your version. The UI Settings tab remains source of truth when JSON and UI disagree.

When should you hire vs DIY this fix?

DIY when you can name every Continue node, none of them sit in front of customer contact or money, and you can spend an afternoon on the staging rungs. Hire when the hide is already in production and you cannot tell which green runs wrote last-valid data.

SituationDIYHire / Audit
Staging only, no customer writesYesNo
Production enrichment, skip log missingMaybe — add the log this weekIf you cannot put a table in front of the pin
Production CRM / email / invoice on ContinueNo — stop the writes todayYes, then reconstruct what green runs did
Never Error on every HTTP “so it never breaks”NoYes
Error workflow unset on activated graphsAttach it todayAudit if you also lack a DLQ
You only tested via Execute WorkflowFinish automatic proofIf you cannot publish a throwaway fail

Spurlock Studios will not quote a fake “percent of workflows we found on Continue.” Open the Settings tab. The dropdown is the evidence.

If you DIY, still use the handbook’s fail-closed default: Stop on writes, Error Trigger for the page, DLQ for the work, Continue only as a labeled skip. If you hire, buy the audit for blast radius, not for someone to sprinkle Continue on Fail until the demo stays green.

A week of DIY that only flips enrichment nodes and ignores write-node Continue is a week you already had. Start at the writes.

FAQ

When does Continue on Fail hide real API errors in n8n?

When the node fails and the execution still succeeds. Error Workflow only runs on a failed execution, so Continue (and Continue using error output, if you ignore the pin) keeps the API error off the pager. On Error → Continue also forwards last valid data, so the hide can include a wrong write, not just a missing alert.

How do I measure whether Continue on Fail is hiding real API errors in n8n?

Open success executions and inspect Continue / Never Error nodes for error objects, empty items, and non-2xx bodies, then count those greens. Failed-execution dashboards cannot see this class of miss. Pair that count with “Error Trigger volume vs write-node volume” — zero pages plus lots of writes is the tell.

What usually fails first when teams try this?

They enable Continue so one bad item will not stop a batch, and last-valid data poisons the next write. Second place: HTTP Never Error with no status IF, so vendor 500 bodies land in the CRM as success. Third: they test the Error Trigger with Manual Execute and conclude the handler is broken when it is working as documented.

How long does this take to show results?

Finding Continue upstream of writes is usually hours on a small instance, not a quarter. Proving Error Trigger attach is one published automatic failure. Reconstructing which green runs already wrote last-valid data can take longer if you never stored inputs — that reconstruction is the cost of the mute, not the cost of the fix.

What should I skip if I only have a week?

Skip polishing enrichment skip-copy. Do not skip: Stop Workflow on every write, attach the Error workflow, turn Never Error off (or IF the status), and run the staging 401 rungs. Leave optional lookups on Continue (error output) only if you can write a skip row in the same week.

When is this not worth doing yet?

When the workflow is unpublished, writes nothing, and has no customer. The moment it is activated on CRM, mail, or money, Continue-as-default is already worth removing. A prototype can fail loud in the editor. Production cannot use green as a substitute for a handler.

CTA

Continue is a branch. Green is not a health check.

Read the Production n8n handbook, then use automation or book the $500 Automation Audit.

FAQ

What questions does this article answer?

When does Continue on Fail hide real API errors in n8n?
When the node fails and the execution still succeeds. Error Workflow only runs on a failed execution, so Continue (and Continue using error output, if you ignore the pin) keeps the API error off the pager. On Error → Continue also forwards last valid data, so the hide can include a wrong write, not just a missing alert.
How do I measure whether Continue on Fail is hiding real API errors in n8n?
Open success executions and inspect Continue / Never Error nodes for error objects, empty items, and non-2xx bodies, then count those greens. Failed-execution dashboards cannot see this class of miss. Pair that count with "Error Trigger volume vs write-node volume" — zero pages plus lots of writes is the tell.
What usually fails first when teams try this?
They enable Continue so one bad item will not stop a batch, and last-valid data poisons the next write. Second place: HTTP Never Error with no status IF, so vendor 500 bodies land in the CRM as success. Third: they test the Error Trigger with Manual Execute and conclude the handler is broken when it is working as documented.
How long does this take to show results?
Finding Continue upstream of writes is usually hours on a small instance, not a quarter. Proving Error Trigger attach is one published automatic failure. Reconstructing which green runs already wrote last-valid data can take longer if you never stored inputs — that reconstruction is the cost of the mute, not the cost of the fix.
What should I skip if I only have a week?
Skip polishing enrichment skip-copy. Do not skip: Stop Workflow on every write, attach the Error workflow, turn Never Error off (or IF the status), and run the staging 401 rungs. Leave optional lookups on Continue (error output) only if you can write a skip row in the same week.
When is this not worth doing yet?
When the workflow is unpublished, writes nothing, and has no customer. The moment it is activated on CRM, mail, or money, Continue-as-default is already worth removing. A prototype can fail loud in the editor. Production cannot use green as a substitute for a handler.
Sources

Last reviewed

More from this lane

Automation

All →
Book the audit