Why did my automation send twice
A second send is usually a webhook retry, a second trigger, a missing key, Wait+resume, or a stalled worker. Match that pair in Executions, then fix it.
William Spurlock Founder — Spurlock Studios 31 MIN
Your automation sent twice because two n8n executions both reached an irreversible node — Gmail, Slack, a CRM create, a Stripe follow-up — and nothing in the graph stopped the second one. The usual classes are a provider webhook retry after a slow or missing 200, two triggers wired to the same send, a missing or late idempotency key, a Wait node that resumed on timeout and again on the callback, or a queue worker that looked stalled after the first run had already sent.
This is the diagnostic. How you install the key is Idempotency Keys in n8n. Do not copy that gate into this page and call it a post-mortem. The production spine those keys sit in is the Production n8n handbook. Across 600+ automations built and 500+ live, the Tuesday ticket is almost never “the Gmail node is haunted.” It is two executions, one side effect, no owner who can pause.
I will not invent a studio-wide “X% of doubles are retries” figure. Those rates only exist after you match your execution pair to your provider event ID.
The short answer
- Two executions, same event ID → provider retry or editor/manual replay. Ack 200 faster. Gate the send.
- Two executions, different event IDs → two triggers, a second workflow, or a human plus a machine both firing.
- One execution, two sends → Wait timeout plus callback, Retry On Fail on the send node, or a loop that hits Gmail twice.
- Queue mode is not a lock. It moves work to workers. It does not make Gmail unique.
- Pause before you replay. The next click is how a diagnostic becomes a third email.
Why did my automation send twice?
A “double send” is a business outcome, not a node color. Two green executions that each emailed the customer are the same incident as one execution that emailed, waited, then emailed again. Classify the pair before you add nodes.
| Class | What you see in Executions | What the customer got | First check |
|---|---|---|---|
| Webhook retry | Two runs, same payload id, seconds to days apart | Two emails / two CRM rows | Webhook Respond mode and provider delivery log |
| Double trigger | Two runs, different payload IDs or different trigger nodes | Two emails from “the same form” | Canvas trigger count; Schedule + Webhook; parent + child |
| Missing key | Two runs, same id, both passed the send node | Two of everything irreversible | No unique claim before Gmail / HubSpot create |
| Wait + resume | One long execution, or one timed-out plus a later resume | Reminder and confirmation, or two confirmations | Wait Limit Wait Time + $execution.resumeUrl retries |
| Queue worker | Two jobs, same execution ID story, or stall then replay | Send landed, then landed again on Retry | EXECUTIONS_MODE=queue, mains, webhook pool, worker logs |
| Retry On Fail / loop | One execution, send node invoked twice | Two emails seconds apart, one row in Executions | Node Settings → Retry On Fail; Split in Batches reset |
If you cannot put the incident in one of those five rows, you do not have a mystery. You have not opened Executions yet.
Decision list:
- Same provider event ID on both runs → retry or missing key (often both).
- Different IDs, same customer → double trigger or two workflows.
- One execution ID, two Gmail successes in the log → Wait, Retry On Fail, or a loop.
- Self-hosted queue, stall language in worker logs → worker class, then still check the key.
- Doubles started the week you published n8n 2.0 → parent/child Wait output change before you blame Redis.
How do you find the duplicate pair in Executions?
Do this before you “just add a filter.” The pair tells you the class. Guessing the class is how you “fix” retries by deleting the Schedule trigger you still needed.
- Pause the workflow. Unpublish it (n8n 2.x) or deactivate it (1.x), or disable the production Webhook. A live graph will keep sending while you stare. n8n 2.0 replaced the active toggle with publish / unpublish. Same kill switch. Different button.
- Open Executions for that workflow. Filter the window around the customer timestamp (their screenshot is usually 1–15 minutes late).
- Pick the two (or more) runs that hit the send node. Ignore runs that stopped at IF / Error.
- Compare trigger node names. Same Webhook twice is a retry candidate. Webhook + Schedule, or Webhook + Execute Workflow, is a double-trigger candidate.
- Diff the inbound JSON. Write down
id/event.id/form_id/hs_object_idfrom each run. Same vs different is the whole fork. - Open each run’s node log for Gmail / Slack / HTTP POST. Note the timestamp of the send, not the trigger. A Wait can separate them by hours.
- Check
execution.retryOfif the payload exposes it, and whether a human hit Retry on the first run after Gmail already returned 200. - Ask the provider. Stripe Event deliveries, HubSpot webhook logs, Typeform delivery log: did they POST twice? Their retry UI is evidence, not folklore.
- Only then pick a class from the table above and apply the matching fix. Do not install three fixes for one pair.
| Evidence | Means | Does not mean |
|---|---|---|
Same evt_… / form response ID | At-least-once delivery or replay | The Gmail node “double-fired by itself” |
| Different IDs, same email address | Two events or two triggers | The provider “duplicated the person” |
| Send timestamps minutes apart, same ID | Retry window or Wait | Instant canvas bug |
| Send timestamps identical to the second | Concurrent workers or two triggers in the same second | You can skip the key |
| First run error, second run success, two emails | First run still sent, then you retried | “The error rolled it back” |
Checklist before you change the graph:
- Workflow is paused
- Two execution URLs are in the ticket
- Event IDs copied, not remembered
- Send-node timestamps recorded
- Provider delivery log opened (or marked “vendor has no UI”)
- Owner named who can unpause (ownership and runbooks)
If step 1 is still published, stop. You are taking notes on a moving fire.
Executions columns lie in predictable ways. Use them, then distrust them.
| Column | Useful for | Lie it tells |
|---|---|---|
| Status | Failed vs success | Success includes “emailed twice, both 200” |
| Started | Ordering the pair | Queue delay makes start ≠ send |
| Trigger | Retry vs second trigger | “Webhook” on both rows still might be two URLs |
Retry / retryOf | Human or n8n retry of a run | Empty when the provider retried with a new execution |
| Duration | Wait vs instant double | A 40-minute Wait looks healthy |
Worked pair (no vendor dashboard required)
You do not need Stripe’s UI to classify the first fork. You need two JSON blobs.
- Execution A, 14:02: Webhook → HubSpot → Gmail. Body
id=evt_abc. Gmail node success at 14:02:08. - Execution B, 14:02:41: same Webhook, same
evt_abc, Gmail success at 14:02:47. - Same ID + two Gmail successes = retry or missing key. Not a second form submit.
- If B’s
idwereevt_defand the email address matched, that is two events (or two triggers). Do not install a retry ack for that.
Write those four lines in the ticket. The rest of this post is only useful after they exist.
How do webhook retries look different from a second trigger?
Providers deliver at least once. Stripe’s webhook docs say endpoints might receive the same event more than once, and live destinations are retried for up to three days with exponential backoff. Sandbox retries three times over a few hours. A dashboard resend does not cancel automatic retries already queued. That is their contract. It is not an n8n defect.
n8n’s Webhook node starts a new execution on every delivery. The Respond setting decides whether the provider is still waiting while you talk to HubSpot.
| Webhook Respond | When the caller gets a status | Retry risk |
|---|---|---|
| Immediately | As soon as n8n accepts (“Workflow got started”) | Low from timeout; you still need a key for the work that continues |
| When Last Node Finishes | After the last node | High if the graph is slower than the provider timeout |
| Using ‘Respond to Webhook’ Node | When that node runs | Low if you respond before HubSpot; high if Respond sits after Gmail |
| Streaming | Chunked | Wrong tool for a fire-and-forget intake |
HubSpot retries failed webhook notifications up to 10 times over 24 hours on connection failure, a timeout longer than five seconds, or any 4xx/5xx. Typeform marks a delivery failed if you take longer than 30 seconds. Hold the HTTP connection through a CRM upsert and you asked for the second POST.
A 500 after a successful send is a retry invitation. The first execution already emailed. The provider only saw failure. They send the same id again. Without a key, Gmail runs again.
| Signal | Retry | Second trigger |
|---|---|---|
Payload id | Identical | Different (or missing) |
| Provider delivery UI | “Retry” / “resend” / 5xx / timeout | Separate subscription, or a second app |
| Trigger node | Same Webhook | Two nodes, or Webhook + Schedule |
| Spacing | Backoff-shaped (seconds → hours → days) | Human-shaped (form submit + “also notify Slack”) |
- Production URL is the one in the vendor dashboard, not the Test URL
- Respond happens before the slow CRM call
- Duplicate path still returns 2xx (see the key post — do not 500 “to be honest”)
- You did not also register the Test URL in the vendor
Test vs production URLs are different registrations. If both are live in the vendor, that is a double trigger wearing a retry costume.
Stripe’s undelivered-event guide is the ack rule in vendor language: if you already processed event.id, skip the work and still return a successful response so automatic retries stop. A 4xx/5xx after Gmail succeeded does the opposite. n8n’s Webhook docs also note that if the workflow errors before the first Respond to Webhook node (when you use that mode), the caller gets 500. Put Respond on the duplicate branch and the success branch. Do not leave the only Respond node after a flaky CRM call.
Procedure when the vendor UI exists:
- Open the delivery for
evt_abc(or the form response ID). - List attempts: timestamps, HTTP status, latency.
- If attempt 1 is 5xx or timeout and attempt 2 is 2xx, you are looking at their retry, not a second customer action.
- If both attempts are 2xx and n8n still has two executions, they delivered twice after you acked — still at-least-once. The key is the remaining lock.
- If there is only one vendor attempt and two n8n executions, stop blaming the provider. Look at double trigger, Wait, or queue.
- Vendor attempt count written next to the two execution IDs
- HTTP statuses copied, not paraphrased
- You know whether attempt 1 was timeout vs 5xx vs 2xx
What does a double trigger look like on the canvas?
A double trigger is two starts that both think they own the send. The payloads can look “the same person” while being two events.
| Pattern | Why it sends twice | What to disable or merge |
|---|---|---|
| Webhook + native HubSpot trigger | Site form POSTs n8n and HubSpot fires contact.creation | Pick one intake. The other becomes an audit log, not a sender |
| Schedule + Webhook | Cron sweeps “unsent” rows the webhook already sent | Cron keys by job + row id + date, or the webhook marks sent first |
| Form Trigger + site Webhook | n8n hosted form and the website’s own POST | One public form |
| Parent Execute Workflow + child’s own Webhook | Parent calls child; vendor also hits the child URL | Child is sub-only, or parent is the only caller |
| Two workflows, same Gmail node copy | “Notify” graph and “CRM” graph both email | One graph after a shared gate, or suffixed keys |
| Editor Execute + live production | You replayed a run that production also received | Never Execute a money/email graph against prod credentials |
n8n will not warn you that two triggers share a Gmail node. The canvas is a graph, not a mutex.
Procedure to hunt extras:
- Count trigger nodes on the published workflow (Webhook, Schedule, Form, app triggers, Chat, Email).
- Search the instance for the same Gmail credential + similar subject line on other workflows.
- Search HubSpot / Stripe / Typeform for two destinations pointing at n8n.
- If a parent uses Execute Workflow, open the child and confirm it has no production Webhook of its own — or that the parent never waits-and-retries the child after the child already sent.
- If you use queue mode with extra webhook processors, confirm the load balancer is not also sending
/webhook/*to a second n8n that still executes locally.
Checklist:
- One intake owns the customer email
- Cron cannot select rows the webhook already sent
- Child workflows are not independently webhook-callable in production
- Vendor dashboard shows one destination URL
- Nobody is clicking Execute on the production workflow “to test”
Two triggers is an architecture mistake. A key still helps — the same person can legitimately submit twice — but it will not merge a form POST and a HubSpot native event that used different IDs. That merge is an upsert rule, not a retry gate.
How do I confirm the key was missing or too late?
If the pair is the same event ID and both runs reached Gmail, you do not have a key in front of the send. Or you have a comment that looks like a key.
This page does not teach the insert. That walkthrough — event ID vs $execution.id, processing vs completed, Stripe Idempotency-Key on outbound POSTs — is Idempotency Keys in n8n. Here you only need to prove the gap.
| Test on the two executions | Missing / late key | Key is doing its job |
|---|---|---|
| Both runs show the Gmail (or create) node as executed | Yes | No — duplicate should have branched |
| Key store has zero row for that event ID | Missing | — |
| Key row written after Gmail timestamp | Late (crash window) | Claim must move before the send |
Key is $execution.id | Useless for retries | Switch to provider event ID |
| Key is a hash of the full body | Often misses (signature / timestamp changed) | Identity fields only |
| Second run 500’d after send | You invited the next retry | 200 on already-complete |
Decision list:
- If there is no store at all (no Postgres table, no Data table, no Redis
SET NX) → you are running at-least-once hope. Install the gate from the key post before you unpause. - If the store exists but the claim sits after Gmail → move it. A timeout after send is the textbook double.
- If the store uses execution ID → it will never see Stripe’s second POST as a duplicate. Change the formula. Do not add a second Gmail IF.
- If the second run no-op’d and the customer still got two emails → you have a double trigger or a second workflow. The key worked. Keep hunting.
n8n Remove Duplicates is a filter, not a lock under two parallel deliveries. If that node is your only gate on a customer email, treat it as missing for this diagnostic.
- You can point at the row that should have blocked run #2
- That row’s
created_atis before the send timestamp - The second execution branched before Gmail
- You have not used
$execution.idas the business key
A missing key is the most common reason the second execution was allowed to send. It is not always the reason there was a second execution. Fix both.
How does Retry On Fail double a send inside one execution?
Inbound retries create two executions. Node Retry On Fail creates two vendor calls inside one execution. The Executions list shows a single green (or red) run. The customer still got two emails.
This is the class people miss because they only compared execution IDs.
| Setting | Safe on | Unsafe on |
|---|---|---|
| Retry On Fail, 3 attempts, wait 1s | GET a status endpoint, list rows, download a file | Gmail send, Slack post, HubSpot create, Stripe POST without Idempotency-Key |
| Continue On Fail | Optional enrichment you will log as skipped | The only path to “we notified the customer” |
| HTTP Request timeout + auto retry | Idempotent GET | POST that creates an object the API will not upsert |
Procedure:
- Open the send node’s Settings. Note Retry On Fail, max tries, wait between tries.
- Open the execution JSON for that node. Multiple attempts on one node look like repeated request metadata, not a second workflow run.
- Compare Gmail’s “sent” timestamps inside the same execution. Two timestamps seconds apart with one execution ID is this class.
- Turn Retry Off on irreversible nodes. If the HTTP node must retry a Stripe POST, the header and the workflow key belong in idempotency keys — not a blind retry toggle.
- For reads that deserve retries, keep a bound (a number you can say out loud) and a backoff. Unlimited retries are how a down API becomes a self-inflicted DDoS.
Loop / Split in Batches is the cousin:
- Split in Batches does not re-queue the same item after a partial fail without a sent flag
- A loop back to Gmail is keyed per item, not “the whole batch again”
- You did not pin an IF to
$itemIndexthat resets on retry
One execution can still be a double send. Count send-node invocations, not rows in the Executions table.
How does Wait + resume create a second send?
The Wait node pauses an execution and, for waits of 65 seconds or more, offloads data to the database, then reloads it on resume. Resume options include a time interval, a specified time, On Webhook Call, and On Form Submitted.
On Webhook Call generates a URL at runtime. n8n exposes $execution.resumeUrl so you can send that URL to a vendor or a human. The URL is unique to that execution. Partial executions change the resume URL — n8n’s own limitation — so the node that emails the URL must run in the same execution as the Wait.
The duplicate patterns that actually email twice:
| Wait setup | What happens | What the customer sees |
|---|---|---|
Limit Wait Time on, plus a late $execution.resumeUrl POST | Timeout path runs the “we didn’t hear back” send; callback later runs the “thanks” send | Two emails for one approval |
| Vendor retries the resume URL | Wait already continued; retry may 404, or a second Wait in a loop consumes it | Two status updates, or a ghost second execution |
| Two Wait nodes, suffix not applied | Docs: $resumeWebhookUrl does not auto-include Webhook Suffix — you append it | Callback hits the wrong Wait, graph continues twice-shaped |
| Wait < 65s, worker/process killed | Short wait stays in-process; a restart plus a new trigger looks like a replay | Duplicate if the first send already left the box |
| Send-and-wait Slack / similar | Queue mode routes these on /webhook-waiting/* | Mis-routed balancer can replay or drop, then the vendor retries |
n8n’s queue-mode load-balancer notes tell you to send /webhook-waiting/* to the webhook processor pool, not to invent a second listener (enable queue mode). If that path hits the wrong process, resume looks haunted.
Procedure when the pair is one long execution:
- Find the Wait node in the log. Note Resume mode and whether Limit Wait Time is on.
- Note the timestamp Wait entered vs the timestamp it left.
- If Limit Wait Time fired, look at the branch after timeout. Does that branch send? If yes, the callback branch must not send the same message.
- Search vendor logs for POSTs to
/webhook-waiting/or the resume path. Two POSTs = retry of the callback. - Confirm the node that distributed
$execution.resumeUrlran in the same execution (no partial Execute that minted a dead URL, then a full run that minted a live one). - Split the messages: timeout = “still waiting” to internal Slack; callback = the only customer-facing send. Or key
approvalId:timeoutvsapprovalId:accepted.
- Timeout branch does not call Gmail with the same copy as the callback branch
- Resume URL is not also registered as a Webhook trigger on another workflow
- Webhook Suffix is actually appended if two Waits exist
-
/webhook-waiting/*goes to webhook processors if you scaled intake
Wait is not a duplicate of “retry.” It is a second clock. Treat it as one.
Parent waits on a child that already sent
n8n 2.0 changed what a parent receives when a child Wait-resumes. Previously, a parent waiting on a child that entered waiting (Wait > 65s, webhook, form, HITL Slack) could get the child’s input back. In 2.0 the parent gets the child’s output after resume (v2.0 breaking changes). If both graphs send “confirmation,” you can get a child send and a parent send for one approval.
| Version | Parent receives after child Wait | Double-send risk |
|---|---|---|
| n8n 1.x | Often the child’s original input | Parent may re-run logic as if the child never worked |
| n8n 2.x | Child’s output after resume | Parent may send again on the output if it still has a Gmail node |
Decision list:
- Child sends the customer email; parent only logs → keep Gmail off the parent.
- Parent sends; child is compute-only → child must not also Gmail.
- Both send today → pick one. The 2.0 output change will not merge them.
If you upgraded through 2.0 and doubles started after the deploy, this is the first Wait question, not Redis.
When do queue workers look like they ran the job twice?
Queue mode is a concurrency split: main (or webhook processors) accept the trigger, Redis holds the job, a worker runs the graph. n8n is explicit: production executions are processed by workers; the HTTP client stays connected to main/webhook; Redis is required; SQLite is not the database for this setup. Share N8N_ENCRYPTION_KEY or workers cannot use credentials.
Queue mode does not claim exactly-once sends. It claims “a worker picked this execution ID.”
| Misconfig / behavior | Why it looks like a double | What n8n actually documents |
|---|---|---|
| Main in the webhook load-balancer pool and separate webhook processors | Two processes can accept the same POST | Do not add main to that pool; optional N8N_DISABLE_PRODUCTION_MAIN_PROCESS=true |
Two n8n start processes without multi-main | Timer / poller class work is supposed to be at-most-once on the leader | Multi-main is Enterprise, needs N8N_MULTI_MAIN_SETUP_ENABLED, sticky sessions |
| Worker lock vs long Wait | QUEUE_WORKER_LOCK_DURATION defaults to 60000 ms; renew every 10000 ms; stall check every 30000 ms (queue env vars) | A worker that cannot renew looks stalled |
| Stalled-job retries on older n8n | Community graphs show jobs re-queued after a stall | n8n 2.0 removed QUEUE_WORKER_MAX_STALLED_COUNT and Bull auto-retry of stalled jobs (breaking changes). Do not assume 2.x will replay the job for you |
| Human Retry after a stall | First worker already sent Gmail, then died before “success” | Second execution is you, not Redis |
| Concurrency flag vs DB pool | Many workers at concurrency 1 can exhaust Postgres connections | n8n recommends worker concurrency 5 or higher |
Hedge, on purpose: I am not going to quote an internal “stalls per week” number. If your version is 2.x, read the changelog — automatic stalled retries are gone. If your version is 1.x, a stall can still re-queue. Either way, a send that already left the worker is not unsent by a crash.
Procedure for a suspected worker double:
- Confirm
EXECUTIONS_MODE=queueon main, workers, and webhook processors. - Count
n8n start/ main containers. More than one without multi-main enabled is a timer suspect. - Confirm the load balancer:
/webhook/*and/webhook-waiting/*→ webhook pool;/webhook-test/*→ main; UI/API → main. - Grep worker logs for stall / lock language around the send timestamp.
- Check whether someone retried the execution in the UI after Gmail’s 200.
- Still install the key. Queue topology fixes extra jobs. It does not make Gmail idempotent.
- One leader for timers (or a single main)
- Main not in the webhook pool
- Encryption key identical everywhere
- You know your n8n major version before you blame Bull
- No UI Retry on a run that already sent
Webhook processors add another fork. n8n lets you scale incoming HTTP with dedicated n8n webhook processes behind a load balancer. If that balancer also includes main, the same POST can be accepted on a process that still executes locally and enqueued for a worker. Their own note: do not put main in the webhook pool; if you disable production webhooks on main (N8N_DISABLE_PRODUCTION_MAIN_PROCESS), keep main out of that pool.
| Process | Should accept /webhook/*? | Should run the graph? |
|---|---|---|
| Main (editor / API) | No, once processors exist | No for production webhooks |
| Webhook processor | Yes | No — it should enqueue |
| Worker | No | Yes |
If you cannot draw that table for your compose file, you do not have a queue topology. You have extra containers.
Queue mode is optional infrastructure. The handbook’s line still holds: switch when a single process is drowning. It is not a reliability upgrade for missing keys.
What should I pause first?
Order matters. Unpause-and-hope is a third send.
- Unpublish / deactivate the workflow (or disable the production Webhook path). This stops new deliveries from becoming new executions. On n8n 2.x the control is Unpublish, not a silent toggle.
- Do not return 5xx from a still-active endpoint “to make it stop.” That extends Stripe’s three-day retry window. If you must keep the URL live, return 200 and no-op.
- Mute only the customer channel (Gmail node → Continue-off, or disconnect the wire) if you need the graph up for CRM upserts. Prefer full pause.
- Tell the owner. The runbook names who can unpause. If nobody is named, you found the other incident — see automation ownership and runbooks.
- Copy the two execution URLs into the ticket before n8n pruning deletes them.
- Check the provider UI for queued retries. Stripe will keep trying unless the destination starts returning 2xx on that
event.id. - Do not click Retry on either execution until you know whether Gmail already succeeded.
| Still active? | Risk | Allowed while paused |
|---|---|---|
| Workflow on | Third send | Nothing customer-facing |
| Workflow off, vendor retries | Vendor noise, no n8n send | Read-only diagnosis |
| Workflow on, 500 | More retries, longer | Never the incident response |
| Workflow on, 200 no-op | Retries drain | Acceptable if you cannot pause |
Pause is cheaper than a clever IF. You can turn it back on after the class is fixed.
How do I fix each class without making a third send?
Match the class you already proved. Mixing fixes is how you hide the next double.
| Class | Fix (this week) | Not the fix |
|---|---|---|
| Webhook retry | Respond immediately or via Respond to Webhook before HubSpot; return 200 on duplicates; then install the key | Adding a 4-second Wait “so they don’t retry” |
| Double trigger | Delete or disconnect the extra trigger; one vendor destination; cron cannot select already-sent rows | A longer Schedule interval |
| Missing / late key | Move or add the atomic claim before Gmail — follow idempotency keys | Remove Duplicates as the money/email gate |
| Wait + resume | Timeout branch ≠ customer send; suffix Wait URLs; route /webhook-waiting/* correctly | Removing Wait and hoping the human is faster |
| Queue worker | One main (or real multi-main); main out of webhook pool; shared encryption key; no UI Retry after send | “Turn off queue mode” as the first move if the pair is a Stripe retry |
Retry On Fail on a Gmail or “create invoice” node is already covered above: n8n re-invokes the node, the vendor created the object on attempt 1, attempt 2 creates another. Bound retries on reads. Do not enable them on irreversible POSTs. If the HTTP node must retry a Stripe POST, send a stable Idempotency-Key — that header pattern lives in the key post.
Split in Batches that restarts the whole set after item 7 failed will re-send items 1–6 unless each item is keyed. Fix with a sent flag on the row, not a smaller batch size.
Procedure after the graph change:
- Keep production paused.
- Clone or use staging credentials.
- POST the same fixture payload twice to the staging Webhook.
- Confirm one customer-facing send, two 200s if it is a retry test.
- If you have Wait, fire timeout without callback, then fire callback, confirm one customer send.
- Only then unpause. Watch the first live duplicate hit (key no-op) in Executions.
- Staging double-POST done
- Wait timeout + callback done if Wait exists
- Error workflow still attached
- Runbook updated with the class you hit
- Unpause owner is the same human who paused
If you cannot run the double-POST, you did not fix it. You edited a canvas.
What is the failure mode after you already emailed twice?
The graph is the easy part. The inbox is not.
Failure mode: the second send already left Gmail. You pause n8n. You add a key. You do not get those emails back. Support now owns a conversation you created.
| Already happened | Do this | Do not do this |
|---|---|---|
| Two identical emails | One human apology; say the second was a retry; do not automate a third “sorry” until that graph is keyed | A new workflow that emails “ignore the last one” on a schedule |
| Two invoices / two charges | Finance reconcile; Stripe dashboard is source of truth; refund path is human | Replay the n8n execution “to reverse it” |
| Two CRM contacts | Merge by email / external ID; keep the one with activity | Delete both and recreate from n8n |
Two Slack #sales pings | Pin the real one; mark the duplicate in-thread | Page the channel with a fourth bot message |
| Wait timeout + approval both sent | Tell the customer which instruction stands | Silence, hoping they pick the later timestamp |
Checklist for the cleanup hour:
- Customer-facing channel handled by a human, once
- CRM merge done before the next sales sequence
- Provider retry queue draining on 200, not still 500ing
- Executions exported or screenshotted before prune
- Runbook has the class name and the date
The cost of that hour is why how much automation costs counts failure cleanup in the bill, not only n8n Cloud. I will not invent your cleanup invoice. Add the time you actually spent.
When should I hire vs DIY this diagnostic?
DIY when you can finish the Executions procedure in this post, name the class, and pause without asking Slack who has the password. Hire when the send is money, legal, or a customer you cannot email a third time while you learn Wait suffixes.
| Situation | Default | Why |
|---|---|---|
| Same event ID, Respond is Last Node Finishes, no key | DIY this afternoon | Classic retry; key post + Respond move |
| Two vendor destinations, two workflows | DIY if you own both canvases | Delete one intake |
Wait + queue + /webhook-waiting + multiple processors | Hire or slow down | Easy to make a third send while “fixing” resume |
| Charges already doubled | Human finance first, then a builder | Do not DIY refunds through Retry |
| No named owner, builder is gone | Ownership work first | The diagnostic will rot in a private Notion |
| You cannot POST twice in staging | You are not ready to unpause | Staging is the test; production is the incident |
Spurlock Studios’ offer on this lane is a call: start at automation or the $500 Automation Audit. That price is the audit, not a promise that your duplicate class is “simple.”
If you only have a week, do the pause, the pair, the Respond-mode fix, and the key in front of Gmail. Skip a queue-mode redesign unless the evidence is two mains or a poisoned webhook pool. Queue work without evidence is a new outage.
DIY also fails when the credential is a personal Gmail OAuth the builder still owns. You can classify the pair and still be unable to pause. That is an ownership bug first — fix the named human and the company-owned credential, then the graph. The audit is for when the class is mixed (Wait + queue + two vendor destinations) and a wrong unpause costs more than $500 of cleanup.
FAQ
Why did my automation send twice?
Because two executions — or one execution with two send paths — both reached an irreversible node. The usual classes are a provider webhook retry, a second trigger, a missing or late idempotency key, a Wait timeout plus callback, or a queue/worker replay after the first send already left. Open Executions, compare event IDs, then fix that class. n8n will not collapse those into one Gmail by default.
How do I measure whether why did my automation send twice is working?
Count unique provider event IDs versus Gmail (or create) node executions in a fixed window, and count key-store hits that branched to no-op. A working fix shows duplicate deliveries in Executions with one customer send and two 2xx responses. If event IDs and sends stay 1:1 while customers still report doubles, you are measuring the wrong workflow or a Wait timeout branch you did not include.
What usually fails first when teams try this?
They skip pause and replay the failed execution after Gmail already returned 200, which is a third send. The next failure is treating Remove Duplicates or $execution.id as a lock, then declaring the graph “safe.” The third is returning 500 after a successful send so the provider’s retry window stays open. Prove the pair, then change one class.
How long does this take to show results?
The pause is immediate: new customer sends stop when the workflow is unpublished. Classifying the pair is usually minutes if Executions and the vendor delivery log still exist. The Respond-mode change stops new timeout-retries on the next delivery. The key only proves itself when the next duplicate POST arrives — run that POST in staging the same day, or you are waiting on production to surprise you.
What should I skip if I only have a week?
Skip a queue-mode rebuild, a multi-main rollout, and a new observability stack unless the pair already points at two mains or a webhook pool that includes main. Do pause, match the pair, move Respond before the slow node, and put a real key in front of Gmail. Leave Wait redesign for graphs that actually use Wait. Leave Stripe refund policy to finance, not to a new node.
When is this not worth doing yet?
If the workflow cannot send, charge, or create a CRM row, a duplicate execution is noise — fix it when you add a send. If you have no staging URL and no way to pause production, diagnosis without a kill switch will create the third send; get ownership and a pause path first. If the “double” is two different customers with similar names, it is not this incident. Spend the week on the graphs that already email humans.
CTA
Two executions. One inbox. Pause, match the pair, then fix that class — not a random new IF.
The spine is in the production n8n handbook. When you want this installed as an operating standard, start at automation or book the $500 Automation Audit.
What questions does this article answer?
- Why did my automation send twice?
- Because two executions — or one execution with two send paths — both reached an irreversible node. The usual classes are a provider webhook retry, a second trigger, a missing or late idempotency key, a Wait timeout plus callback, or a queue/worker replay after the first send already left. Open Executions, compare event IDs, then fix that class. n8n will not collapse those into one Gmail by default.
- How do I measure whether why did my automation send twice is working?
- Count unique provider event IDs versus Gmail (or create) node executions in a fixed window, and count key-store hits that branched to no-op. A working fix shows duplicate deliveries in Executions with **one** customer send and two 2xx responses. If event IDs and sends stay 1:1 while customers still report doubles, you are measuring the wrong workflow or a Wait timeout branch you did not include.
- What usually fails first when teams try this?
- They skip pause and replay the failed execution after Gmail already returned 200, which is a third send. The next failure is treating Remove Duplicates or `$execution.id` as a lock, then declaring the graph "safe." The third is returning 500 after a successful send so the provider's retry window stays open. Prove the pair, then change one class.
- How long does this take to show results?
- The pause is immediate: new customer sends stop when the workflow is unpublished. Classifying the pair is usually minutes if Executions and the vendor delivery log still exist. The Respond-mode change stops *new* timeout-retries on the next delivery. The key only proves itself when the **next** duplicate POST arrives — run that POST in staging the same day, or you are waiting on production to surprise you.
- What should I skip if I only have a week?
- Skip a queue-mode rebuild, a multi-main rollout, and a new observability stack unless the pair already points at two mains or a webhook pool that includes main. Do pause, match the pair, move Respond before the slow node, and put a real key in front of Gmail. Leave Wait redesign for graphs that actually use Wait. Leave Stripe refund policy to finance, not to a new node.
- When is this not worth doing yet?
- If the workflow cannot send, charge, or create a CRM row, a duplicate execution is noise — fix it when you add a send. If you have no staging URL and no way to pause production, diagnosis without a kill switch will create the third send; get ownership and a pause path first. If the "double" is two different customers with similar names, it is not this incident. Spend the week on the graphs that already email humans.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation Why is one giant n8n canvas a production liability
One giant n8n canvas is a production liability because debug, credentials, retries, and deploys share one blast radius. Split it with named contracts.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.