When Automation Fails at 2am: Alerts, Severity, and Who Gets Woken
Overnight failures are auth, schema, and silent success. Page money and customer paths, morning-triage the rest, and heartbeat never-ran gaps before dawn.
William Spurlock Founder — Spurlock Studios Updated 22 MIN
When automation fails at 2am, the failure is almost never a red node. It is an expired grant, a payload that no longer matches the write, or a 2xx that wrote nothing. One of three things should happen: a human gets woken for irreversible work, a morning queue gets a ticket for everything else, or a heartbeat proves the trigger never fired. Platform defaults do none of that well.
Spurlock Studios builds for the overnight case first. Happy-path demos lie. The three overnight modes — auth, schema, silent success — sit on top of the production spine in the Production n8n handbook. This post owns who wakes, what waits, and how you detect a green lie before dawn.
The short answer
- Page only when money moves, a customer gets contacted, or a system of record goes wrong with no safe retry.
- Morning-triage enrichment skips, rate-limit waits, and noncritical sync lag.
- Detect silence with heartbeats. Error alerts fire when something ran and failed. They miss never-ran and silent-success.
- Name an owner before go-live. “The founder might see Slack” is not on-call.
- Zapier / Make / n8n all fail quietly until you add severity, schema asserts, and auth pause yourself.
What actually fails at 2am?
Across 500+ automations, the overnight pile-up is not random. It clusters into three modes. The rail looks fine for two of them.
| Mode | What the rail shows | What the business sees at 9am | Typical overnight cause |
|---|---|---|---|
| Auth | 401 / invalid_grant / Zap or scenario auto-off | CRM went dark; no new rows | Access token died; refresh token revoked or unused |
| Schema | Green run, or a mapping warning nobody pages | Wrong field, null in a required write, type flip | Vendor renamed a property; null arrived where a string used to |
| Silent success | 2xx / “success” / empty array | Invoice not marked paid; lead never created | HTTP success with no side effect; filter dropped the batch |
Auth is the only mode that reliably throws. Schema and silent success often complete. That is why “we have error email” is not overnight coverage.
Decision list — classify the 2am event before you page:
- Did a node throw, or did the run finish green?
- If it threw, is it auth (
401,invalid_grant) or a transient503/ rate limit? - If it finished green, did the system of record actually change?
- If nothing ran, is the heartbeat late?
Throwing is the easy case. Green-and-wrong is the expensive one.
Why do platforms stay quiet by default?
Out of the box, most rails treat failure as a UI badge or a polite email to the account owner. That email often lands in a shared inbox nobody checks at night. The workflow may keep “running” while every item fails, or — worse — stop receiving events and look healthy because nothing errored.
| Default behavior | What operators think | What actually happens |
|---|---|---|
| Error email to account owner | Someone is on call | Inbox mute or spam folder |
Slack webhook to #ops | Humans will wake | Channel muted after week one |
| Red execution in the UI | Visible overnight | Visible only if someone opens the app |
| Zap / scenario auto-off | Safe stop | Silent stop; backlog grows |
| Custom error handler “handled” | Alert still fires | Often it does not |
Zapier will automatically turn a Zap off when it errors at least 95% of the time across more than 20 runs in seven days. Team plans get a 24-hour grace email; Enterprise gets 72 hours. After that, the Zap is off. Nothing incoming is an error. It is a hole.
Make emails on unhandled errors and when repeated failures disable the schedule. Turn on incomplete executions and those errors become warnings. Warnings do not disable the schedule and do not count toward the consecutive-error cutoff. The scenario keeps running. The broken bundles sit in a tab you only open if you remember to.
n8n’s Error Trigger only runs on automatic executions, not manual tests. If the trigger node itself fails, you may get no execution.id and no execution URL. Attach the error workflow in settings or you have a red row and nobody to tell.
Quiet failure is the product default. Loud, graded failure is something you design.
How does overnight auth failure look?
Access tokens are short-lived on purpose. Google’s own docs: user access tokens expire after one hour. The refresh token is supposed to mint a new one while nobody is watching. Overnight is when that mint fails.
RFC 6749 §5.2 names the grant failure invalid_grant: the authorization grant or refresh token is invalid, expired, revoked, or was issued to another client. Google lists the usual reasons a refresh token stops working: user revoke, six months unused, Gmail scopes after a password change, a 100-token-per-client cap that silently invalidates the oldest token, time-boxed access, admin policy, or a consent screen still in Testing (seven-day refresh lifetime for external users).
| Auth signal | Meaning overnight | Page? |
|---|---|---|
| HTTP 401 on a write | Access token dead; refresh may still work | P2 unless money path already looping |
invalid_grant on refresh | Refresh is dead. Human must re-consent | P1 if the path writes CRM / billing; else P2 at 08:00 |
| Zap / scenario auto-disabled after error ratio | Auth (or anything) crossed the vendor cutoff | P1 if the path is customer-facing |
| 403 on an API that uses 403 for expiry | Refresh never triggered if the rail only watches 401 | Treat as auth until proven otherwise |
n8n refreshes OAuth reactively when the status matches the expired-token code (default 401). A vendor that returns 403, or 200 with {"error":"invalid_token"} in the body, will not refresh. The run can look like a normal HTTP failure — or, if you swallow the body, like a success. The credential playbook is OAuth tokens that stop expiring quietly. This page owns the page/pause decision when that refresh dies at 2am.
What breaks: a founder’s personal Google login, unused for a week of vacation. Refresh dies. Nightly CRM sync 401s. Error email hits an inbox on Do Not Disturb. Morning: a day of missed leads.
What you do instead:
- Prefer a service account or workspace app over a personal OAuth grant.
- Treat
invalid_grantas pause-class. Stop the workflow. Do not retry into a dead grant. - Page if the path is money or customer-contact. Otherwise ticket for 08:00 with “re-consent required.”
- Write
last_auth_ok_atnext tolast_success_at. Auth silence is a different heartbeat.
- Production credentials are not a founder’s Gmail
-
invalid_grantroutes to pause, not to a retry storm - Auth failures include the credential name, not only the node name
- Re-consent steps live in the runbook next to the phone number
What is a silent-success overnight run?
Silent success is a finished execution whose HTTP status is fine and whose business outcome is missing. The rail will not page you. It did its job: it got a 2xx.
RFC 9110 defines 204 No Content as success with no payload body. A 200 with [] or {"updated":0} is also success. Automation graphs that treat “status < 300” as “the CRM row exists” are lying to themselves.
Stripe makes the other half of this trap explicit. Their webhook docs tell you to return a 2xx quickly — before the accounting write — and they retry live deliveries for up to three days if you do not. If you ack 200 and then the worker dies, Stripe is satisfied. Your ledger is not. That is silent success until you reconcile event IDs against applied writes.
| Silent-success shape | Rail status | Business status | How you catch it |
|---|---|---|---|
204 / empty 200 on a write you expected to echo | Success | Unknown | Assert a returned id or a follow-up read |
200 + {"matched":0} / empty array | Success | No row touched | Fail the node when matched count is 0 |
| Filter / IF dropped every item | Success (nothing to do) | Batch vanished | Alert when inbound count > 0 and outbound count = 0 |
| Zapier custom error handler ran | Handled — no error email | Step failed, path “succeeded” | Handler must page; do not rely on Zapier mail |
| Make incomplete execution stored | Warning, schedule still on | Bundle parked | Checkpoint the incomplete-execution queue, not the error inbox |
Zapier’s own error-handler article is blunt: when a custom handler runs, Zapier will not send error notification emails. If your “coverage” is the default mail, the handler just turned it off.
Procedure — prove a write, not a status:
- Capture inbound count and a business key (email, invoice id, Stripe
event.id). - After the write, require one of: returned record id,
matched >= 1, or a read-back. - If the assert fails, Stop and Error (or the rail equivalent) so the error workflow fires.
- Store
last_success_atonly after the assert, never after the HTTP node.
A green execution that skipped the assert is a demo. Demos do not get overnight trust.
How does schema drift land after midnight?
Vendors ship field changes on their schedule. Yours is 2am because that is when the cron runs. JSON Schema is explicit about why this is quiet: properties listed under properties are optional unless named in required, extra keys are allowed unless you close the object, and a property set to null is present — it is type null, not “missing.” That is all in the object keyword reference.
Overnight schema failures look like this:
| Drift | Payload | What a loose mapping does | What a closed schema does |
|---|---|---|---|
| Field rename | email → emailAddress | Writes null / empty to CRM email | Rejects; error workflow fires |
| Type flip | "100" instead of 100 | Concatenates, or the API coerces wrong | Rejects type |
| Null in required | "email": null | CRM accepts empty email | Fails type: string and required |
| Extra field you now depend on | New plan_tier ignored | Downstream default = free | additionalProperties: false or an explicit allow-list |
| Envelope change | { data: { ... } } instead of flat | Every mapped path is undefined | Rejects; no write |
Make already classifies a slice of this as bundle / data / incomplete-data errors. Those are useful when they throw. They are useless when the module accepts the bundle and writes a blank.
What breaks: a nightly invoice sync. The billing API wraps the old object in data. n8n still maps amount from the root. Amount is undefined. The HTTP node posts { "amount": null }. The vendor returns 200 and stores zero. Accounting is wrong until someone compares totals.
What you do instead:
- Validate the payload against a versioned schema before any irreversible write.
- Put money fields, emails, and external ids in
required. - Fail on
nullfor those fields. Missing and null are different; both are poison for a write. - On failure, dead-letter the payload and page if the path is P1. Do not “default to zero.”
- Schema file (or n8n validation node) is pinned to a version, not “whatever the last sample was”
- Required list includes every field a 9am human would sue over
- Staging replayed a renamed-field fixture and a null-email fixture
- Error alert includes the failed keyword (
required,type), not only “validation failed”
Schema is a trust boundary. If you validate only in the editor’s sample payload, you validated January.
When should a human get paged?
Write severity before you wire Slack. Copy this table into the runbook. PagerDuty maps critical / error to high urgency (escalates) and warning / info to low urgency (does not escalate). Missing severity defaults to high. If you dump every n8n error in as critical, you will train the team to mute the pager.
| Severity | Overnight examples | Response |
|---|---|---|
| P1 — wake someone | Payment capture failed mid-charge; CRM overwrite; outbound SMS/email blast; invalid_grant on a billing credential; silent-success on an invoice write | Phone / PagerDuty / SMS within minutes |
| P2 — morning first | Lead sync delayed; enrichment API down; schema reject on a noncritical enrich; auth fail on a reporting read | Ticket + owner Slack by start of business |
| P3 — backlog | Optional research step skipped; soft validation warning | Weekly triage board |
Rule: if the blast radius can create refunds, legal risk, or a customer-facing lie before 9am, it is P1. Auth on a money path is P1. Schema reject that blocked a write is P1 if waiting makes the damage worse. Silent success on an invoice is P1 because you will not see a red node.
Page when all of these are true:
- The side effect is irreversible or customer-visible
- Waiting until morning makes the damage worse (duplicates, wrong quotes, missed SLAs)
- A human action in the next hour can stop or reverse it
Do not page for:
- A node that already retried and will retry again safely
- Enrichment that is allowed to fail open
- Staging / test workflows
- Known vendor maintenance windows you already documented
- A 401 that refresh already healed on the same execution
If every failure pages, people mute the channel. Mute is how 2am incidents become 9am discoveries.
How do you detect a workflow that never ran?
Error workflows answer “this execution failed.” They do not answer “the webhook died,” “the cron never fired,” or “every run was a silent success.” n8n is explicit: the Error Trigger receives details when a linked workflow fails. A workflow that does not start sends nothing. A workflow that succeeds with an empty write also sends nothing.
Add a dead-man / heartbeat check:
- Every asserted production success writes
last_success_atto a small store (DB row, Airtable, Redis key). - A separate schedule (every 15–60 minutes, matched to expected volume) checks that timestamp.
- If
now - last_success_atexceeds the SLA for that flow, fire a silence alert with severity based on the path. - Optionally write
last_run_atas well.last_run_atfresh andlast_success_atstale is the silent-success signature.
| Trigger type | Silence signal | Typical SLA to alert |
|---|---|---|
| High-volume webhook | No asserted success for N minutes during business hours | 15–30 min |
| Nightly cron | Missed expected window | Window end + 30 min |
| Weekly report | Missed Monday 06:00 | +2 hours |
| Auth-sensitive path | last_auth_ok_at older than 2× token lifetime | 2 hours after expiry window |
Silence detection is the control most “Slack alert” tutorials skip. Pair it with the credential lifecycle in the OAuth expiry post so a dead grant and a dead trigger do not look the same in the page.
How do Zapier, Make, and n8n differ overnight?
Same ops problem; different knobs. None of them wake a human until you decide the message and the destination.
| Rail | Common overnight default | Overnight hole | What you must add |
|---|---|---|---|
| Zapier | Error email; Zap may turn off after the 95% / 20-run rule | Auto-off looks like “no errors”; custom handlers suppress mail | Routed alerts, owner, silence check, write asserts |
| Make | Email on unhandled errors; auto-disable after consecutive errors | Incomplete executions turn errors into warnings | Watch the incomplete queue; do not treat Warning as healthy |
| n8n | Error workflow if you attach one | Manual tests never fire it; trigger-node failures omit execution URL | Alert contract, heartbeat, schema assert, auth pause |
n8n wins when you want one handler attached to every production flow. It does not wake anyone until you decide what the message says and who receives it. The handbook is the rest of the spine. This post is the overnight contract on top of it.
Checklist — rail-specific overnight proof:
- Zapier: error-ratio override is a conscious choice, not an accident; handlers page themselves
- Make: incomplete-execution count is on a schedule, not a tab you remember
- n8n: every production workflow has an error workflow set; you proved it with an activated (not manual) failure
- All three: heartbeat is a separate workflow, not a hope that volume continues
What belongs in the overnight ownership contract?
Before activation, fill this once. If a line is “TBD,” the workflow is not production.
Workflow: _______________
Owner (primary): _______________
Backup owner: _______________
P1 channel: _______________
P2 channel: _______________
Mute policy: no mute on P1; P2 may snooze until 08:00 local
Heartbeat key: _______________
Max silence: _______________
Auth pause steps: _______________
Schema version: _______________
Write assert: _______________
Rollback / pause steps: _______________
If the primary is on vacation and the backup is “TBD,” you have a demo with a schedule. Demos do not get customer data after midnight.
- Primary and backup are humans with phone numbers, not role names
- Backup has credential access before the first P1
- P1 channel cannot be muted by the same people who get enrichment noise
- Pause steps are written at the same time as the Slack webhook
What happens when the alerts channel gets muted?
What breaks: a chatty Error Workflow posts every rate-limit hiccup into #alerts. After three nights, the team mutes the channel. On night four, a payment path fails — or worse, succeeds with matched: 0 and never posts — and nobody sees it.
What it costs: morning discovery, manual cleanup, trust hit with whoever owns the CRM.
What you do instead:
- Split channels:
#automation-p1(never mute) and#automation-triage(morning). - Route by severity inside the error handler — do not post everything once.
- Cap repeats: after N identical errors in an hour, collapse to one “still failing” message with a count.
- Keep P1 on a pager tool if chat culture cannot protect the channel.
- Do not put silent-success asserts and enrichment skips on the same destination.
| Noise source | Put it here | Never put it here |
|---|---|---|
invalid_grant on billing | P1 pager | #alerts dump |
| Schema reject on invoice | P1 pager | Email-to-founder |
| Enrichment 429 | Triage | Pager |
| Heartbeat miss on nightly cron | P1 if money path; else triage at window+30 | Nowhere (this is the usual miss) |
| Handled Zapier error | Handler’s own page | Assumption that Zapier mailed you |
Mute is a symptom of undifferentiated severity. Fix the routing. Do not ask people to “check Slack more.”
What does morning triage actually mean?
Morning triage is not “ignore until angry.” It is a named queue with an SLA:
- Owner opens
#automation-triagebefore first customer calls - Sort by customer-visible impact, not by timestamp
- Pause anything still failing in a loop — especially auth
- Replay or discard parked items (Make incomplete executions, n8n DLQ) with a written reason
- File one changelog note if a vendor caused schema drift
- Confirm heartbeats flipped back to green after the first good asserted run
If morning triage regularly spills past noon, you undersized ownership or over-automated enrichment noise into the wrong bucket.
| Morning item | First action | Done looks like |
|---|---|---|
| Auth P2 | Re-consent or rotate; do not replay until grant is live | One successful asserted run |
| Schema P2 | Diff last good payload vs last reject; pin a new schema version | Fixture replayed in staging, then prod |
| Silent-success suspect | Compare inbound count vs write count for the night | Gap closed or items dead-lettered |
| Auto-disabled Zap / scenario | Read the error ratio; fix the cause; only then re-enable | Heartbeat SLA reset |
A triage that only “acks Slack” is a status LED. Status LEDs do not restore data.
Who has pause authority after hours?
Paging without pause rights creates spectators. The on-call person must be able to:
- Deactivate the workflow (or flip a feature flag / dry-run)
- Rotate or disconnect a bad credential
- Tell sales/support the sync is paused
- Open the parked-item queue and stop replaying poison
- Leave the schema pin in place so a bad deploy cannot silently reopen the write
Write those steps in the runbook next to the phone number. An alert that only says “failed” without pause authority is a status LED.
Overnight pause order we actually use:
- Disable the trigger or flip dry-run. Side effects stop first.
- Snapshot the last payloads and the last error. Do not “just retry” from memory.
- If auth: revoke/rotate, then one staging call, then prod.
- If schema: pin the reject, do not widen the schema at 3am to “make it green.”
- If silent success: stop acking 2xx until the write assert is in. A wider 200 is how you lose the next night too.
Bravery is not a restore strategy. Pause is.
Overnight checklist (before you call it production)
- Severity table exists for this workflow
- P1 has a phone/SMS path, not only Slack
- P2 has a morning owner named in writing
- Heartbeat / dead-man check covers “never ran”
-
last_success_atis written only after a write assert - Auth failures pause; they do not retry into
invalid_grant - Schema is versioned; required money/email/id fields cannot be null
- Error handler includes workflow name, execution link, failed node, severity, mode (auth / schema / silent / throw)
- Mute policy documented for the P1 channel
- Pause steps written (who flips the workflow off)
- Staging proved one intentional failure for an activated path
- Staging proved one silent-success fixture (
200+ empty write) and one renamed-field fixture
If a box is unchecked, you have a weekday demo. Do not put customer data on its cron.
Decision list: page or wait
Ask in order:
- Can waiting until morning create irreversible customer or money damage? → Page
- Is the failure “never ran” rather than “ran and failed”? → Silence alert (P1 if the path is P1)
- Did it run green with no asserted write? → Silent-success (P1 on money / customer-contact; else P2)
- Is it
invalid_grantor a refresh that will not recover without a human? → Pause, then page or ticket by path - Is it a known transient with bounded retry still in budget? → Wait; log
- Is it enrichment / optional enrichment? → Morning triage
- Unsure? → Treat as P1 once, then downgrade with evidence
Unsure defaults to loud once. Habitual over-paging defaults to mute. Calibrate with real incidents, not vibes.
Alert template worth pasting
Use one shape across Zapier digests, Make notifications, and n8n Error Workflows:
[P1] {{workflowName}} failed
Mode: auth | schema | silent-success | throw | silence
Node: {{failedNode}}
Error: {{errorMessage}}
Exec: {{executionUrl}}
Owner: {{ownerPrimary}} (backup {{ownerBackup}})
Next: {{nextAction}}
Assert?: {{writeAssertPassed}}
Last success: {{lastSuccessAt}}
For silence alerts, Mode: silence and Last success carry the story. For silent-success, Assert?: no is the line that tells the human this was not a red node. Same channel discipline, different signal.
n8n will give you workflow name, last node, and an execution URL when the execution was saved — see the Error Trigger payload. It will not give you the inbound body. Pull that from your own log or a get-execution call if you need to replay at 2am.
How this fits the spine
Overnight posture sits next to idempotency, parked-item queues, and schema checks — not instead of them. Alerts without a write assert create panic about the wrong night, or silence about the right one. Alerts without ownership create noise. The handbook is the full list. This post owns who wakes, why, and how auth / schema / silent success present after midnight.
If you only “email the account owner,” you do not have overnight coverage. You have hope.
FAQ
Is a Zapier error email enough overnight?
No. Account-owner email is not an on-call system, and Zapier stops mailing when a custom error handler runs. Route P1 to a human who answers overnight, keep P2 for morning, and add silence detection so a Zap that quietly turned off still surfaces.
Why does OAuth fail at 2am without a useful alert?
Access tokens expire on a timer — Google’s are typically one hour — and refresh can return invalid_grant while you sleep. The rail may auto-disable after an error ratio, or swallow a 403 that never triggers refresh. Error email then hits a muted inbox. Pause on grant death, page money paths, and treat credentials as a first-class overnight control.
What is a silent-success overnight failure?
A finished run with a 2xx (or a “handled” status) that did not apply the business write. Empty arrays, matched: 0, dropped filters, and a Stripe-style 200-before-work all qualify. The rail will not page you. Assert the write, store last_success_at only after that assert, and heartbeat the gap.
How do I catch schema drift after midnight?
Validate required fields and types before the write. JSON Schema treats listed properties as optional until you mark them required, and null is a value, not an absence. Fail loud on rename, type flip, and null-in-required, then park the payload. Page if the path can lie to a customer before 9am.
What belongs in an n8n Error Workflow alert?
Minimum: workflow name, execution URL, failed node, error message, severity, named owner, failure mode (auth / schema / silent / throw / silence), and whether the write assert ran. n8n’s Error Trigger only fires on automatic failures. Without those fields, the alert is noise people learn to ignore.
Who is the named owner after hours?
A real person (and a backup) with authority to pause the workflow and access to credentials. “The agency” or “whoever built it” is not a name. Write primary and backup before activation, and give the backup pause rights before the first P1 — not during it.
CTA
If your automations only “email the account owner,” you do not have overnight coverage — you have hope.
For a production review of severity, heartbeats, and the three overnight modes, start at automation or book a call.
What questions does this article answer?
- Is a Zapier error email enough overnight?
- No. Account-owner email is not an on-call system, and Zapier [stops mailing](https://help.zapier.com/hc/en-us/articles/22495436062605-Set-up-custom-error-handling) when a custom error handler runs. Route P1 to a human who answers overnight, keep P2 for morning, and add silence detection so a Zap that quietly turned off still surfaces.
- Why does OAuth fail at 2am without a useful alert?
- Access tokens expire on a timer — Google's are typically one hour — and refresh can return `invalid_grant` while you sleep. The rail may auto-disable after an error ratio, or swallow a 403 that never triggers refresh. Error email then hits a muted inbox. Pause on grant death, page money paths, and treat credentials as a first-class overnight control.
- What is a silent-success overnight failure?
- A finished run with a 2xx (or a "handled" status) that did not apply the business write. Empty arrays, `matched: 0`, dropped filters, and a Stripe-style 200-before-work all qualify. The rail will not page you. Assert the write, store `last_success_at` only after that assert, and heartbeat the gap.
- How do I catch schema drift after midnight?
- Validate required fields and types before the write. JSON Schema treats listed properties as optional until you mark them `required`, and `null` is a value, not an absence. Fail loud on rename, type flip, and null-in-required, then park the payload. Page if the path can lie to a customer before 9am.
- What belongs in an n8n Error Workflow alert?
- Minimum: workflow name, execution URL, failed node, error message, severity, named owner, failure mode (auth / schema / silent / throw / silence), and whether the write assert ran. n8n's Error Trigger only fires on automatic failures. Without those fields, the alert is noise people learn to ignore.
- Who is the named owner after hours?
- A real person (and a backup) with authority to pause the workflow and access to credentials. "The agency" or "whoever built it" is not a name. Write primary and backup before activation, and give the backup pause rights before the first P1 — not during it.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation Why doesn’t worker concurrency cap my n8n sub-workflows
Worker concurrency does not cap n8n sub-workflows. Each Execute Workflow child is a new execution the production limit skips, usually on the parent worker.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.