Why Your Automation Broke on a Tuesday
Six hundred workflows in, the failures are boringly predictable. Four structures prevent almost all of them, and none of them are exciting.
William Spurlock Founder — Spurlock Studios Updated 12 MIN
Nobody’s automation fails during the demo. It fails eleven weeks later, on a Tuesday, when a third-party API starts returning null in a field that has always been a string, and your pipeline cheerfully writes eight hundred empty records into a CRM that a salesperson is about to open.
I have built roughly six hundred of these. The failures are not creative. They are the same four failures, and they have the same four fixes. The operating model for putting n8n behind those fixes is the Production n8n handbook. How we scope a first production rail is on /automation. This post is the Tuesday post-mortem, written before Tuesday.
The short answer
- Webhooks are delivered at least once. Gate irreversible work on an idempotency key computed from the event, not from the run.
- Do not retry a half-finished multi-step job. Park it on a dead-letter queue with the input, the execution id, and a replay link.
- Validate the shape of every external payload before it touches your database. Pause on mismatch.
- The first version of anything that spends money, emails a customer, or deletes a record proposes and waits.
- “It retries” is not an error path. An error path is a queue a human can open.
Why did the same webhook run twice?
Webhooks are delivered at least once. Not exactly once. Every provider you integrate with will, eventually, deliver the same event twice — usually because their first delivery attempt timed out on your end after you had already processed it.
If that event charges a card, sends an email, or increments a counter, you now have a support ticket. The pipeline did what you asked, twice. The provider did what their docs said. Your assumption was the part that was wrong.
The fix is an idempotency key computed from the payload itself, checked against a store before anything irreversible happens:
const key = createHash("sha256")
.update(`${payload.id}:${payload.updated_at}`)
.digest("hex");
if (await seen.has(key)) return { status: 200, note: "duplicate" };
await seen.set(key, true, { ttl: 60 * 60 * 24 * 7 });
Two nodes. It prevents the entire class.
A few details that make the two nodes actually work. Key the business event, not the workflow execution. Execution ids are unique per delivery, which is the opposite of what you want. Prefer the provider’s event id when they give you one; hash a stable subset of the payload when they do not. Claim the key before the charge, the email, or the CRM create. Storing it after the write is how a timeout still doubles you: the write succeeded, the ack failed, the retry arrives, the key is missing, and you do it again.
Return 200 on the duplicate. A 500 after a successful write invites the next retry. The provider is not trying to hurt you. They are trying to get an ack. Give them the ack. Keep the side effect singular.
The full key-and-claim pattern, including why you should not use the n8n execution id as the key, is in idempotency keys in n8n. Here the Tuesday lesson is smaller: if you cannot point at the store that says “we already did this event,” you do not have exactly-once effects. You have at-least-once hope.
Duplicate checklist:
- Key comes from the payload (event id or a stable hash), never from the run id
- Key is claimed before the irreversible node
- Duplicate deliveries still return 2xx
- TTL is longer than the provider’s retry window
- You have fired the same payload twice in staging and seen one side effect
When does a retry make the failure worse?
The default instinct is to wrap the failing step in a retry. This is correct roughly half the time, and actively harmful the other half, because retrying a partially-applied multi-step operation re-applies the steps that already succeeded.
A timeout on “create invoice” after “create customer” is not a reason to start the graph over. Starting over is how you get two customers and one angry finance thread. A 503 from a downstream API is a reason to retry that step, with a bound, with backoff. The difference is whether the failure is a symptom of the network or a symptom of the work.
Failures belong in a dead-letter queue, not a retry loop. Route the exception out of the main thread with three things attached: the original input, the execution ID, and the error. Then notify a human with a one-click replay link. You keep the data, you keep the ability to fix and re-run, and you stop the pipeline from thrashing against an API that is down.
Unbounded retry on a poison payload is a load test against someone else’s API, billed to you. Bound the retry. Then park. Replay from the failed step with the same idempotency key so the parts that already succeeded stay succeeded.
The parking-lot version of this argument is dead letter queues for automations. The wake-up version is an error workflow operators actually read. Slack is the pager. The queue is the work. If the error only exists as a red execution in a UI nobody opens until a customer complains, you do not have an error path. You have a diary.
Retry versus park:
- Retry: transient 429/503, timeout with no side effect yet, a credential that is about to succeed after a refresh you already automated.
- Park: schema mismatch, partial write, auth that will not succeed until a human rotates it, anything you do not understand.
- Never auto-redrive a payload that failed validation. You will re-poison the same destination.
What happens when the schema changes under you?
This is the Tuesday failure. An upstream provider ships a change, a field goes from string to null, and because most automation tools are permissive by default, the bad value propagates all the way to your database.
Permissive is a demo feature. In production it means “we will type whatever arrived.” Empty company name. Null email. A record that looks created and is useless. The salesperson opens the CRM and the pipeline looks healthy because every node returned 200. HTTP 200 is not a contract.
Put a validator immediately after every external call. Not a big one — a shape check:
const Contact = z.object({
id: z.string(),
email: z.string().email(),
company: z.string().min(1),
});
When validation fails, the item goes to the review queue and the run pauses. A paused pipeline is a five-minute inconvenience. A poisoned database is a weekend.
Pause is the part teams skip. They log the error and keep going, which is how eight hundred empty records get a head start. Fail closed on required fields. Fail closed on type flips. Fail closed on empty strings that used to be identifiers. Then patch the mapping against a fresh payload in staging, then replay the parked items.
The longer drift playbook is when APIs change, automations break: pin versions when the vendor offers them, run the same assertions against recorded live samples, and treat changelog mail as decoration until a named owner turns it into calendar work. Tuesday is what happens when none of that exists and the vendor shipped anyway.
Schema checklist:
- Every external payload hits a shape check before a write
- Required fields are required in the validator, not in a comment
- Mismatch pauses the irreversible path, not just logs
- Parked items keep the original JSON
- Replay happens after the mapping is fixed, not while it is still wrong
Why was the workflow allowed to do something irreversible?
The last failure is a design failure rather than an engineering one: the workflow was allowed to do something irreversible without anyone agreeing to it.
My rule is that anything which spends money, contacts a customer, or deletes a record gets a human gate by default. Not forever — you move the gate once the numbers earn it. But the first version of every pipeline proposes and waits.
“Proposes and waits” is a state, not a vibe. The workflow writes a draft invoice, a draft email, a draft CRM merge. A human clicks. Then the irreversible node runs, under the same idempotency key you would have used anyway. If you cannot show me the approval step, you shipped autonomy you did not earn.
Autonomy is a promotion, not a default. You earn it with a watch window of understood errors, not with a clean demo week. Demo week does not include the null field, the double webhook, or the intern who replayed an old execution. Production does.
If the business wants the gate gone on day one, they are buying a faster way to send the wrong thing. You can still build that. You should still name it.
Gate checklist:
- Money, customer contact, and deletes start behind a human
- The proposal is inspectable (the actual email body, the actual amount)
- Approval is a recorded action, not a Slack thumbs-up that cannot be replayed
- Promotion off the gate has a named owner and a watch window
- Kill switch is a pause, documented, that a backup owner can hit
What should you ask a builder about their error path?
None of this is interesting. It is four structures, they add maybe fifteen percent to the build, and they are the entire difference between an automation you trust and one you check every morning.
If you are evaluating someone to build these for you, ask them to describe their error path. If the answer is “it retries,” keep looking.
A real answer names the store for idempotency keys, the table that holds dead letters, the validator that can pause a run, and the gate in front of irreversible nodes. It names an owner who gets the alert. It names how you replay without doubling. It sounds slightly boring. Boring is the point.
The ownership half of that conversation — who gets paged when the builder is on a plane — is automation ownership and runbooks. Structures without a named human are still a Tuesday, just a quieter one.
You do not need a more clever workflow. You need the four structures, in the graph, before the first real payload arrives. After it arrives is how you get the weekend.
Why do automations fail on a Tuesday instead of in the demo?
Because the demo used a happy payload, once, with a human watching. Tuesday is when the provider retries a webhook you already handled, a field arrives as null, a retry repeats a write that already succeeded, or a graph sends mail nobody approved. Those are production shapes. They are also the four you can build for in advance.
Is retry-on-fail an error-handling strategy?
Only for transient failures with no side effect yet, and only with a bound. Retrying a half-applied graph, or a payload that failed a shape check, makes the incident larger. Park the work on a dead-letter queue with the original input and a replay path. “It retries” is how pipelines thrash and databases quietly fill with empty rows.
What questions does this article answer?
- Why do automations fail on a Tuesday instead of in the demo?
- Because the demo used a happy payload, once, with a human watching. Tuesday is when the provider retries a webhook you already handled, a field arrives as null, a retry repeats a write that already succeeded, or a graph sends mail nobody approved. Those are production shapes. They are also the four you can build for in advance.
- Is retry-on-fail an error-handling strategy?
- Only for transient failures with no side effect yet, and only with a bound. Retrying a half-applied graph, or a payload that failed a shape check, makes the incident larger. Park the work on a dead-letter queue with the original input and a replay path. "It retries" is how pipelines thrash and databases quietly fill with empty rows.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation Why doesn’t worker concurrency cap my n8n sub-workflows
Worker concurrency does not cap n8n sub-workflows. Each Execute Workflow child is a new execution the production limit skips, usually on the parent worker.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.