Spurlock Studios
Contact
Share LinkedIn X
Amber node beads on a dark rail. Thesis: DID AUTOMATION SEND TWICE.

Your automation sent twice because two n8n executions both reached an irreversible node — Gmail, Slack, a CRM create, a Stripe follow-up — and nothing in the graph stopped the second one. The usual classes are a provider webhook retry after a slow or missing 200, two triggers wired to the same send, a missing or late idempotency key, a Wait node that resumed on timeout and again on the callback, or a queue worker that looked stalled after the first run had already sent.

This is the diagnostic. How you install the key is Idempotency Keys in n8n. Do not copy that gate into this page and call it a post-mortem. The production spine those keys sit in is the Production n8n handbook. Across 600+ automations built and 500+ live, the Tuesday ticket is almost never “the Gmail node is haunted.” It is two executions, one side effect, no owner who can pause.

I will not invent a studio-wide “X% of doubles are retries” figure. Those rates only exist after you match your execution pair to your provider event ID.

The short answer

  • Two executions, same event ID → provider retry or editor/manual replay. Ack 200 faster. Gate the send.
  • Two executions, different event IDs → two triggers, a second workflow, or a human plus a machine both firing.
  • One execution, two sends → Wait timeout plus callback, Retry On Fail on the send node, or a loop that hits Gmail twice.
  • Queue mode is not a lock. It moves work to workers. It does not make Gmail unique.
  • Pause before you replay. The next click is how a diagnostic becomes a third email.

Why did my automation send twice?

A “double send” is a business outcome, not a node color. Two green executions that each emailed the customer are the same incident as one execution that emailed, waited, then emailed again. Classify the pair before you add nodes.

ClassWhat you see in ExecutionsWhat the customer gotFirst check
Webhook retryTwo runs, same payload id, seconds to days apartTwo emails / two CRM rowsWebhook Respond mode and provider delivery log
Double triggerTwo runs, different payload IDs or different trigger nodesTwo emails from “the same form”Canvas trigger count; Schedule + Webhook; parent + child
Missing keyTwo runs, same id, both passed the send nodeTwo of everything irreversibleNo unique claim before Gmail / HubSpot create
Wait + resumeOne long execution, or one timed-out plus a later resumeReminder and confirmation, or two confirmationsWait Limit Wait Time + $execution.resumeUrl retries
Queue workerTwo jobs, same execution ID story, or stall then replaySend landed, then landed again on RetryEXECUTIONS_MODE=queue, mains, webhook pool, worker logs
Retry On Fail / loopOne execution, send node invoked twiceTwo emails seconds apart, one row in ExecutionsNode Settings → Retry On Fail; Split in Batches reset

If you cannot put the incident in one of those five rows, you do not have a mystery. You have not opened Executions yet.

Decision list:

  1. Same provider event ID on both runs → retry or missing key (often both).
  2. Different IDs, same customer → double trigger or two workflows.
  3. One execution ID, two Gmail successes in the log → Wait, Retry On Fail, or a loop.
  4. Self-hosted queue, stall language in worker logs → worker class, then still check the key.
  5. Doubles started the week you published n8n 2.0 → parent/child Wait output change before you blame Redis.

How do you find the duplicate pair in Executions?

Do this before you “just add a filter.” The pair tells you the class. Guessing the class is how you “fix” retries by deleting the Schedule trigger you still needed.

  1. Pause the workflow. Unpublish it (n8n 2.x) or deactivate it (1.x), or disable the production Webhook. A live graph will keep sending while you stare. n8n 2.0 replaced the active toggle with publish / unpublish. Same kill switch. Different button.
  2. Open Executions for that workflow. Filter the window around the customer timestamp (their screenshot is usually 1–15 minutes late).
  3. Pick the two (or more) runs that hit the send node. Ignore runs that stopped at IF / Error.
  4. Compare trigger node names. Same Webhook twice is a retry candidate. Webhook + Schedule, or Webhook + Execute Workflow, is a double-trigger candidate.
  5. Diff the inbound JSON. Write down id / event.id / form_id / hs_object_id from each run. Same vs different is the whole fork.
  6. Open each run’s node log for Gmail / Slack / HTTP POST. Note the timestamp of the send, not the trigger. A Wait can separate them by hours.
  7. Check execution.retryOf if the payload exposes it, and whether a human hit Retry on the first run after Gmail already returned 200.
  8. Ask the provider. Stripe Event deliveries, HubSpot webhook logs, Typeform delivery log: did they POST twice? Their retry UI is evidence, not folklore.
  9. Only then pick a class from the table above and apply the matching fix. Do not install three fixes for one pair.
EvidenceMeansDoes not mean
Same evt_… / form response IDAt-least-once delivery or replayThe Gmail node “double-fired by itself”
Different IDs, same email addressTwo events or two triggersThe provider “duplicated the person”
Send timestamps minutes apart, same IDRetry window or WaitInstant canvas bug
Send timestamps identical to the secondConcurrent workers or two triggers in the same secondYou can skip the key
First run error, second run success, two emailsFirst run still sent, then you retried“The error rolled it back”

Checklist before you change the graph:

  • Workflow is paused
  • Two execution URLs are in the ticket
  • Event IDs copied, not remembered
  • Send-node timestamps recorded
  • Provider delivery log opened (or marked “vendor has no UI”)
  • Owner named who can unpause (ownership and runbooks)

If step 1 is still published, stop. You are taking notes on a moving fire.

Executions columns lie in predictable ways. Use them, then distrust them.

ColumnUseful forLie it tells
StatusFailed vs successSuccess includes “emailed twice, both 200”
StartedOrdering the pairQueue delay makes start ≠ send
TriggerRetry vs second trigger“Webhook” on both rows still might be two URLs
Retry / retryOfHuman or n8n retry of a runEmpty when the provider retried with a new execution
DurationWait vs instant doubleA 40-minute Wait looks healthy

Worked pair (no vendor dashboard required)

You do not need Stripe’s UI to classify the first fork. You need two JSON blobs.

  1. Execution A, 14:02: Webhook → HubSpot → Gmail. Body id = evt_abc. Gmail node success at 14:02:08.
  2. Execution B, 14:02:41: same Webhook, same evt_abc, Gmail success at 14:02:47.
  3. Same ID + two Gmail successes = retry or missing key. Not a second form submit.
  4. If B’s id were evt_def and the email address matched, that is two events (or two triggers). Do not install a retry ack for that.

Write those four lines in the ticket. The rest of this post is only useful after they exist.

How do webhook retries look different from a second trigger?

Providers deliver at least once. Stripe’s webhook docs say endpoints might receive the same event more than once, and live destinations are retried for up to three days with exponential backoff. Sandbox retries three times over a few hours. A dashboard resend does not cancel automatic retries already queued. That is their contract. It is not an n8n defect.

n8n’s Webhook node starts a new execution on every delivery. The Respond setting decides whether the provider is still waiting while you talk to HubSpot.

Webhook RespondWhen the caller gets a statusRetry risk
ImmediatelyAs soon as n8n accepts (“Workflow got started”)Low from timeout; you still need a key for the work that continues
When Last Node FinishesAfter the last nodeHigh if the graph is slower than the provider timeout
Using ‘Respond to Webhook’ NodeWhen that node runsLow if you respond before HubSpot; high if Respond sits after Gmail
StreamingChunkedWrong tool for a fire-and-forget intake

HubSpot retries failed webhook notifications up to 10 times over 24 hours on connection failure, a timeout longer than five seconds, or any 4xx/5xx. Typeform marks a delivery failed if you take longer than 30 seconds. Hold the HTTP connection through a CRM upsert and you asked for the second POST.

A 500 after a successful send is a retry invitation. The first execution already emailed. The provider only saw failure. They send the same id again. Without a key, Gmail runs again.

SignalRetrySecond trigger
Payload idIdenticalDifferent (or missing)
Provider delivery UI“Retry” / “resend” / 5xx / timeoutSeparate subscription, or a second app
Trigger nodeSame WebhookTwo nodes, or Webhook + Schedule
SpacingBackoff-shaped (seconds → hours → days)Human-shaped (form submit + “also notify Slack”)
  • Production URL is the one in the vendor dashboard, not the Test URL
  • Respond happens before the slow CRM call
  • Duplicate path still returns 2xx (see the key post — do not 500 “to be honest”)
  • You did not also register the Test URL in the vendor

Test vs production URLs are different registrations. If both are live in the vendor, that is a double trigger wearing a retry costume.

Stripe’s undelivered-event guide is the ack rule in vendor language: if you already processed event.id, skip the work and still return a successful response so automatic retries stop. A 4xx/5xx after Gmail succeeded does the opposite. n8n’s Webhook docs also note that if the workflow errors before the first Respond to Webhook node (when you use that mode), the caller gets 500. Put Respond on the duplicate branch and the success branch. Do not leave the only Respond node after a flaky CRM call.

Procedure when the vendor UI exists:

  1. Open the delivery for evt_abc (or the form response ID).
  2. List attempts: timestamps, HTTP status, latency.
  3. If attempt 1 is 5xx or timeout and attempt 2 is 2xx, you are looking at their retry, not a second customer action.
  4. If both attempts are 2xx and n8n still has two executions, they delivered twice after you acked — still at-least-once. The key is the remaining lock.
  5. If there is only one vendor attempt and two n8n executions, stop blaming the provider. Look at double trigger, Wait, or queue.
  • Vendor attempt count written next to the two execution IDs
  • HTTP statuses copied, not paraphrased
  • You know whether attempt 1 was timeout vs 5xx vs 2xx

What does a double trigger look like on the canvas?

A double trigger is two starts that both think they own the send. The payloads can look “the same person” while being two events.

PatternWhy it sends twiceWhat to disable or merge
Webhook + native HubSpot triggerSite form POSTs n8n and HubSpot fires contact.creationPick one intake. The other becomes an audit log, not a sender
Schedule + WebhookCron sweeps “unsent” rows the webhook already sentCron keys by job + row id + date, or the webhook marks sent first
Form Trigger + site Webhookn8n hosted form and the website’s own POSTOne public form
Parent Execute Workflow + child’s own WebhookParent calls child; vendor also hits the child URLChild is sub-only, or parent is the only caller
Two workflows, same Gmail node copy“Notify” graph and “CRM” graph both emailOne graph after a shared gate, or suffixed keys
Editor Execute + live productionYou replayed a run that production also receivedNever Execute a money/email graph against prod credentials

n8n will not warn you that two triggers share a Gmail node. The canvas is a graph, not a mutex.

Procedure to hunt extras:

  1. Count trigger nodes on the published workflow (Webhook, Schedule, Form, app triggers, Chat, Email).
  2. Search the instance for the same Gmail credential + similar subject line on other workflows.
  3. Search HubSpot / Stripe / Typeform for two destinations pointing at n8n.
  4. If a parent uses Execute Workflow, open the child and confirm it has no production Webhook of its own — or that the parent never waits-and-retries the child after the child already sent.
  5. If you use queue mode with extra webhook processors, confirm the load balancer is not also sending /webhook/* to a second n8n that still executes locally.

Checklist:

  • One intake owns the customer email
  • Cron cannot select rows the webhook already sent
  • Child workflows are not independently webhook-callable in production
  • Vendor dashboard shows one destination URL
  • Nobody is clicking Execute on the production workflow “to test”

Two triggers is an architecture mistake. A key still helps — the same person can legitimately submit twice — but it will not merge a form POST and a HubSpot native event that used different IDs. That merge is an upsert rule, not a retry gate.

How do I confirm the key was missing or too late?

If the pair is the same event ID and both runs reached Gmail, you do not have a key in front of the send. Or you have a comment that looks like a key.

This page does not teach the insert. That walkthrough — event ID vs $execution.id, processing vs completed, Stripe Idempotency-Key on outbound POSTs — is Idempotency Keys in n8n. Here you only need to prove the gap.

Test on the two executionsMissing / late keyKey is doing its job
Both runs show the Gmail (or create) node as executedYesNo — duplicate should have branched
Key store has zero row for that event IDMissing—
Key row written after Gmail timestampLate (crash window)Claim must move before the send
Key is $execution.idUseless for retriesSwitch to provider event ID
Key is a hash of the full bodyOften misses (signature / timestamp changed)Identity fields only
Second run 500’d after sendYou invited the next retry200 on already-complete

Decision list:

  1. If there is no store at all (no Postgres table, no Data table, no Redis SET NX) → you are running at-least-once hope. Install the gate from the key post before you unpause.
  2. If the store exists but the claim sits after Gmail → move it. A timeout after send is the textbook double.
  3. If the store uses execution ID → it will never see Stripe’s second POST as a duplicate. Change the formula. Do not add a second Gmail IF.
  4. If the second run no-op’d and the customer still got two emails → you have a double trigger or a second workflow. The key worked. Keep hunting.

n8n Remove Duplicates is a filter, not a lock under two parallel deliveries. If that node is your only gate on a customer email, treat it as missing for this diagnostic.

  • You can point at the row that should have blocked run #2
  • That row’s created_at is before the send timestamp
  • The second execution branched before Gmail
  • You have not used $execution.id as the business key

A missing key is the most common reason the second execution was allowed to send. It is not always the reason there was a second execution. Fix both.

How does Retry On Fail double a send inside one execution?

Inbound retries create two executions. Node Retry On Fail creates two vendor calls inside one execution. The Executions list shows a single green (or red) run. The customer still got two emails.

This is the class people miss because they only compared execution IDs.

SettingSafe onUnsafe on
Retry On Fail, 3 attempts, wait 1sGET a status endpoint, list rows, download a fileGmail send, Slack post, HubSpot create, Stripe POST without Idempotency-Key
Continue On FailOptional enrichment you will log as skippedThe only path to “we notified the customer”
HTTP Request timeout + auto retryIdempotent GETPOST that creates an object the API will not upsert

Procedure:

  1. Open the send node’s Settings. Note Retry On Fail, max tries, wait between tries.
  2. Open the execution JSON for that node. Multiple attempts on one node look like repeated request metadata, not a second workflow run.
  3. Compare Gmail’s “sent” timestamps inside the same execution. Two timestamps seconds apart with one execution ID is this class.
  4. Turn Retry Off on irreversible nodes. If the HTTP node must retry a Stripe POST, the header and the workflow key belong in idempotency keys — not a blind retry toggle.
  5. For reads that deserve retries, keep a bound (a number you can say out loud) and a backoff. Unlimited retries are how a down API becomes a self-inflicted DDoS.

Loop / Split in Batches is the cousin:

  • Split in Batches does not re-queue the same item after a partial fail without a sent flag
  • A loop back to Gmail is keyed per item, not “the whole batch again”
  • You did not pin an IF to $itemIndex that resets on retry

One execution can still be a double send. Count send-node invocations, not rows in the Executions table.

How does Wait + resume create a second send?

The Wait node pauses an execution and, for waits of 65 seconds or more, offloads data to the database, then reloads it on resume. Resume options include a time interval, a specified time, On Webhook Call, and On Form Submitted.

On Webhook Call generates a URL at runtime. n8n exposes $execution.resumeUrl so you can send that URL to a vendor or a human. The URL is unique to that execution. Partial executions change the resume URL — n8n’s own limitation — so the node that emails the URL must run in the same execution as the Wait.

The duplicate patterns that actually email twice:

Wait setupWhat happensWhat the customer sees
Limit Wait Time on, plus a late $execution.resumeUrl POSTTimeout path runs the “we didn’t hear back” send; callback later runs the “thanks” sendTwo emails for one approval
Vendor retries the resume URLWait already continued; retry may 404, or a second Wait in a loop consumes itTwo status updates, or a ghost second execution
Two Wait nodes, suffix not appliedDocs: $resumeWebhookUrl does not auto-include Webhook Suffix — you append itCallback hits the wrong Wait, graph continues twice-shaped
Wait < 65s, worker/process killedShort wait stays in-process; a restart plus a new trigger looks like a replayDuplicate if the first send already left the box
Send-and-wait Slack / similarQueue mode routes these on /webhook-waiting/*Mis-routed balancer can replay or drop, then the vendor retries

n8n’s queue-mode load-balancer notes tell you to send /webhook-waiting/* to the webhook processor pool, not to invent a second listener (enable queue mode). If that path hits the wrong process, resume looks haunted.

Procedure when the pair is one long execution:

  1. Find the Wait node in the log. Note Resume mode and whether Limit Wait Time is on.
  2. Note the timestamp Wait entered vs the timestamp it left.
  3. If Limit Wait Time fired, look at the branch after timeout. Does that branch send? If yes, the callback branch must not send the same message.
  4. Search vendor logs for POSTs to /webhook-waiting/ or the resume path. Two POSTs = retry of the callback.
  5. Confirm the node that distributed $execution.resumeUrl ran in the same execution (no partial Execute that minted a dead URL, then a full run that minted a live one).
  6. Split the messages: timeout = “still waiting” to internal Slack; callback = the only customer-facing send. Or key approvalId:timeout vs approvalId:accepted.
  • Timeout branch does not call Gmail with the same copy as the callback branch
  • Resume URL is not also registered as a Webhook trigger on another workflow
  • Webhook Suffix is actually appended if two Waits exist
  • /webhook-waiting/* goes to webhook processors if you scaled intake

Wait is not a duplicate of “retry.” It is a second clock. Treat it as one.

Parent waits on a child that already sent

n8n 2.0 changed what a parent receives when a child Wait-resumes. Previously, a parent waiting on a child that entered waiting (Wait > 65s, webhook, form, HITL Slack) could get the child’s input back. In 2.0 the parent gets the child’s output after resume (v2.0 breaking changes). If both graphs send “confirmation,” you can get a child send and a parent send for one approval.

VersionParent receives after child WaitDouble-send risk
n8n 1.xOften the child’s original inputParent may re-run logic as if the child never worked
n8n 2.xChild’s output after resumeParent may send again on the output if it still has a Gmail node

Decision list:

  1. Child sends the customer email; parent only logs → keep Gmail off the parent.
  2. Parent sends; child is compute-only → child must not also Gmail.
  3. Both send today → pick one. The 2.0 output change will not merge them.

If you upgraded through 2.0 and doubles started after the deploy, this is the first Wait question, not Redis.

When do queue workers look like they ran the job twice?

Queue mode is a concurrency split: main (or webhook processors) accept the trigger, Redis holds the job, a worker runs the graph. n8n is explicit: production executions are processed by workers; the HTTP client stays connected to main/webhook; Redis is required; SQLite is not the database for this setup. Share N8N_ENCRYPTION_KEY or workers cannot use credentials.

Queue mode does not claim exactly-once sends. It claims “a worker picked this execution ID.”

Misconfig / behaviorWhy it looks like a doubleWhat n8n actually documents
Main in the webhook load-balancer pool and separate webhook processorsTwo processes can accept the same POSTDo not add main to that pool; optional N8N_DISABLE_PRODUCTION_MAIN_PROCESS=true
Two n8n start processes without multi-mainTimer / poller class work is supposed to be at-most-once on the leaderMulti-main is Enterprise, needs N8N_MULTI_MAIN_SETUP_ENABLED, sticky sessions
Worker lock vs long WaitQUEUE_WORKER_LOCK_DURATION defaults to 60000 ms; renew every 10000 ms; stall check every 30000 ms (queue env vars)A worker that cannot renew looks stalled
Stalled-job retries on older n8nCommunity graphs show jobs re-queued after a stalln8n 2.0 removed QUEUE_WORKER_MAX_STALLED_COUNT and Bull auto-retry of stalled jobs (breaking changes). Do not assume 2.x will replay the job for you
Human Retry after a stallFirst worker already sent Gmail, then died before “success”Second execution is you, not Redis
Concurrency flag vs DB poolMany workers at concurrency 1 can exhaust Postgres connectionsn8n recommends worker concurrency 5 or higher

Hedge, on purpose: I am not going to quote an internal “stalls per week” number. If your version is 2.x, read the changelog — automatic stalled retries are gone. If your version is 1.x, a stall can still re-queue. Either way, a send that already left the worker is not unsent by a crash.

Procedure for a suspected worker double:

  1. Confirm EXECUTIONS_MODE=queue on main, workers, and webhook processors.
  2. Count n8n start / main containers. More than one without multi-main enabled is a timer suspect.
  3. Confirm the load balancer: /webhook/* and /webhook-waiting/* → webhook pool; /webhook-test/* → main; UI/API → main.
  4. Grep worker logs for stall / lock language around the send timestamp.
  5. Check whether someone retried the execution in the UI after Gmail’s 200.
  6. Still install the key. Queue topology fixes extra jobs. It does not make Gmail idempotent.
  • One leader for timers (or a single main)
  • Main not in the webhook pool
  • Encryption key identical everywhere
  • You know your n8n major version before you blame Bull
  • No UI Retry on a run that already sent

Webhook processors add another fork. n8n lets you scale incoming HTTP with dedicated n8n webhook processes behind a load balancer. If that balancer also includes main, the same POST can be accepted on a process that still executes locally and enqueued for a worker. Their own note: do not put main in the webhook pool; if you disable production webhooks on main (N8N_DISABLE_PRODUCTION_MAIN_PROCESS), keep main out of that pool.

ProcessShould accept /webhook/*?Should run the graph?
Main (editor / API)No, once processors existNo for production webhooks
Webhook processorYesNo — it should enqueue
WorkerNoYes

If you cannot draw that table for your compose file, you do not have a queue topology. You have extra containers.

Queue mode is optional infrastructure. The handbook’s line still holds: switch when a single process is drowning. It is not a reliability upgrade for missing keys.

What should I pause first?

Order matters. Unpause-and-hope is a third send.

  1. Unpublish / deactivate the workflow (or disable the production Webhook path). This stops new deliveries from becoming new executions. On n8n 2.x the control is Unpublish, not a silent toggle.
  2. Do not return 5xx from a still-active endpoint “to make it stop.” That extends Stripe’s three-day retry window. If you must keep the URL live, return 200 and no-op.
  3. Mute only the customer channel (Gmail node → Continue-off, or disconnect the wire) if you need the graph up for CRM upserts. Prefer full pause.
  4. Tell the owner. The runbook names who can unpause. If nobody is named, you found the other incident — see automation ownership and runbooks.
  5. Copy the two execution URLs into the ticket before n8n pruning deletes them.
  6. Check the provider UI for queued retries. Stripe will keep trying unless the destination starts returning 2xx on that event.id.
  7. Do not click Retry on either execution until you know whether Gmail already succeeded.
Still active?RiskAllowed while paused
Workflow onThird sendNothing customer-facing
Workflow off, vendor retriesVendor noise, no n8n sendRead-only diagnosis
Workflow on, 500More retries, longerNever the incident response
Workflow on, 200 no-opRetries drainAcceptable if you cannot pause

Pause is cheaper than a clever IF. You can turn it back on after the class is fixed.

How do I fix each class without making a third send?

Match the class you already proved. Mixing fixes is how you hide the next double.

ClassFix (this week)Not the fix
Webhook retryRespond immediately or via Respond to Webhook before HubSpot; return 200 on duplicates; then install the keyAdding a 4-second Wait “so they don’t retry”
Double triggerDelete or disconnect the extra trigger; one vendor destination; cron cannot select already-sent rowsA longer Schedule interval
Missing / late keyMove or add the atomic claim before Gmail — follow idempotency keysRemove Duplicates as the money/email gate
Wait + resumeTimeout branch ≠ customer send; suffix Wait URLs; route /webhook-waiting/* correctlyRemoving Wait and hoping the human is faster
Queue workerOne main (or real multi-main); main out of webhook pool; shared encryption key; no UI Retry after send“Turn off queue mode” as the first move if the pair is a Stripe retry

Retry On Fail on a Gmail or “create invoice” node is already covered above: n8n re-invokes the node, the vendor created the object on attempt 1, attempt 2 creates another. Bound retries on reads. Do not enable them on irreversible POSTs. If the HTTP node must retry a Stripe POST, send a stable Idempotency-Key — that header pattern lives in the key post.

Split in Batches that restarts the whole set after item 7 failed will re-send items 1–6 unless each item is keyed. Fix with a sent flag on the row, not a smaller batch size.

Procedure after the graph change:

  1. Keep production paused.
  2. Clone or use staging credentials.
  3. POST the same fixture payload twice to the staging Webhook.
  4. Confirm one customer-facing send, two 200s if it is a retry test.
  5. If you have Wait, fire timeout without callback, then fire callback, confirm one customer send.
  6. Only then unpause. Watch the first live duplicate hit (key no-op) in Executions.
  • Staging double-POST done
  • Wait timeout + callback done if Wait exists
  • Error workflow still attached
  • Runbook updated with the class you hit
  • Unpause owner is the same human who paused

If you cannot run the double-POST, you did not fix it. You edited a canvas.

What is the failure mode after you already emailed twice?

The graph is the easy part. The inbox is not.

Failure mode: the second send already left Gmail. You pause n8n. You add a key. You do not get those emails back. Support now owns a conversation you created.

Already happenedDo thisDo not do this
Two identical emailsOne human apology; say the second was a retry; do not automate a third “sorry” until that graph is keyedA new workflow that emails “ignore the last one” on a schedule
Two invoices / two chargesFinance reconcile; Stripe dashboard is source of truth; refund path is humanReplay the n8n execution “to reverse it”
Two CRM contactsMerge by email / external ID; keep the one with activityDelete both and recreate from n8n
Two Slack #sales pingsPin the real one; mark the duplicate in-threadPage the channel with a fourth bot message
Wait timeout + approval both sentTell the customer which instruction standsSilence, hoping they pick the later timestamp

Checklist for the cleanup hour:

  • Customer-facing channel handled by a human, once
  • CRM merge done before the next sales sequence
  • Provider retry queue draining on 200, not still 500ing
  • Executions exported or screenshotted before prune
  • Runbook has the class name and the date

The cost of that hour is why how much automation costs counts failure cleanup in the bill, not only n8n Cloud. I will not invent your cleanup invoice. Add the time you actually spent.

When should I hire vs DIY this diagnostic?

DIY when you can finish the Executions procedure in this post, name the class, and pause without asking Slack who has the password. Hire when the send is money, legal, or a customer you cannot email a third time while you learn Wait suffixes.

SituationDefaultWhy
Same event ID, Respond is Last Node Finishes, no keyDIY this afternoonClassic retry; key post + Respond move
Two vendor destinations, two workflowsDIY if you own both canvasesDelete one intake
Wait + queue + /webhook-waiting + multiple processorsHire or slow downEasy to make a third send while “fixing” resume
Charges already doubledHuman finance first, then a builderDo not DIY refunds through Retry
No named owner, builder is goneOwnership work firstThe diagnostic will rot in a private Notion
You cannot POST twice in stagingYou are not ready to unpauseStaging is the test; production is the incident

Spurlock Studios’ offer on this lane is a call: start at automation or the $500 Automation Audit. That price is the audit, not a promise that your duplicate class is “simple.”

If you only have a week, do the pause, the pair, the Respond-mode fix, and the key in front of Gmail. Skip a queue-mode redesign unless the evidence is two mains or a poisoned webhook pool. Queue work without evidence is a new outage.

DIY also fails when the credential is a personal Gmail OAuth the builder still owns. You can classify the pair and still be unable to pause. That is an ownership bug first — fix the named human and the company-owned credential, then the graph. The audit is for when the class is mixed (Wait + queue + two vendor destinations) and a wrong unpause costs more than $500 of cleanup.

FAQ

Why did my automation send twice?

Because two executions — or one execution with two send paths — both reached an irreversible node. The usual classes are a provider webhook retry, a second trigger, a missing or late idempotency key, a Wait timeout plus callback, or a queue/worker replay after the first send already left. Open Executions, compare event IDs, then fix that class. n8n will not collapse those into one Gmail by default.

How do I measure whether why did my automation send twice is working?

Count unique provider event IDs versus Gmail (or create) node executions in a fixed window, and count key-store hits that branched to no-op. A working fix shows duplicate deliveries in Executions with one customer send and two 2xx responses. If event IDs and sends stay 1:1 while customers still report doubles, you are measuring the wrong workflow or a Wait timeout branch you did not include.

What usually fails first when teams try this?

They skip pause and replay the failed execution after Gmail already returned 200, which is a third send. The next failure is treating Remove Duplicates or $execution.id as a lock, then declaring the graph “safe.” The third is returning 500 after a successful send so the provider’s retry window stays open. Prove the pair, then change one class.

How long does this take to show results?

The pause is immediate: new customer sends stop when the workflow is unpublished. Classifying the pair is usually minutes if Executions and the vendor delivery log still exist. The Respond-mode change stops new timeout-retries on the next delivery. The key only proves itself when the next duplicate POST arrives — run that POST in staging the same day, or you are waiting on production to surprise you.

What should I skip if I only have a week?

Skip a queue-mode rebuild, a multi-main rollout, and a new observability stack unless the pair already points at two mains or a webhook pool that includes main. Do pause, match the pair, move Respond before the slow node, and put a real key in front of Gmail. Leave Wait redesign for graphs that actually use Wait. Leave Stripe refund policy to finance, not to a new node.

When is this not worth doing yet?

If the workflow cannot send, charge, or create a CRM row, a duplicate execution is noise — fix it when you add a send. If you have no staging URL and no way to pause production, diagnosis without a kill switch will create the third send; get ownership and a pause path first. If the “double” is two different customers with similar names, it is not this incident. Spend the week on the graphs that already email humans.

CTA

Two executions. One inbox. Pause, match the pair, then fix that class — not a random new IF.

The spine is in the production n8n handbook. When you want this installed as an operating standard, start at automation or book the $500 Automation Audit.

FAQ

What questions does this article answer?

Why did my automation send twice?
Because two executions — or one execution with two send paths — both reached an irreversible node. The usual classes are a provider webhook retry, a second trigger, a missing or late idempotency key, a Wait timeout plus callback, or a queue/worker replay after the first send already left. Open Executions, compare event IDs, then fix that class. n8n will not collapse those into one Gmail by default.
How do I measure whether why did my automation send twice is working?
Count unique provider event IDs versus Gmail (or create) node executions in a fixed window, and count key-store hits that branched to no-op. A working fix shows duplicate deliveries in Executions with **one** customer send and two 2xx responses. If event IDs and sends stay 1:1 while customers still report doubles, you are measuring the wrong workflow or a Wait timeout branch you did not include.
What usually fails first when teams try this?
They skip pause and replay the failed execution after Gmail already returned 200, which is a third send. The next failure is treating Remove Duplicates or `$execution.id` as a lock, then declaring the graph "safe." The third is returning 500 after a successful send so the provider's retry window stays open. Prove the pair, then change one class.
How long does this take to show results?
The pause is immediate: new customer sends stop when the workflow is unpublished. Classifying the pair is usually minutes if Executions and the vendor delivery log still exist. The Respond-mode change stops *new* timeout-retries on the next delivery. The key only proves itself when the **next** duplicate POST arrives — run that POST in staging the same day, or you are waiting on production to surprise you.
What should I skip if I only have a week?
Skip a queue-mode rebuild, a multi-main rollout, and a new observability stack unless the pair already points at two mains or a webhook pool that includes main. Do pause, match the pair, move Respond before the slow node, and put a real key in front of Gmail. Leave Wait redesign for graphs that actually use Wait. Leave Stripe refund policy to finance, not to a new node.
When is this not worth doing yet?
If the workflow cannot send, charge, or create a CRM row, a duplicate execution is noise — fix it when you add a send. If you have no staging URL and no way to pause production, diagnosis without a kill switch will create the third send; get ownership and a pause path first. If the "double" is two different customers with similar names, it is not this incident. Spend the week on the graphs that already email humans.
Sources

Last reviewed

More from this lane

Automation

All →
Book the audit