How do I manage and maintain AI automations once they're running
You maintain live AI automations with a named owner, a current runbook, a shared error workflow, a weekly review, and pinned model IDs — not a green canvas.
William Spurlock Founder — Spurlock Studios 29 MIN
You manage and maintain AI automations once they’re running by naming an owner who can pause them, keeping a half-page runbook current, attaching a shared error workflow that pages a human, running a weekly review, and pinning model IDs so the vendor cannot swap behavior under you. A green canvas is not a maintenance program. Across 600+ automations built and 500+ live, plus 20,000+ hours on agentic systems, the rails that stay trusted are the ones with a Tuesday loop — not a launch checklist nobody opened after week one.
This spoke sits under the Production n8n handbook. Exit-week ownership and credential cutover live in automation ownership and runbooks. This page is the weekday operating system after nobody is leaving: owner, runbook, error workflow, weekly review, pinned models.
The short answer
- Owner = the person who gets the alert and can pause the workflow today. A Slack channel is not an owner.
- Runbook = half a page that still matches production: purpose, trigger, irreversible steps, pause path, credential names, model pin, last test.
- Error workflow = one shared Error Trigger handler attached under Settings → Error workflow on every live graph.
- Weekly review = failures, silent zero-runs, pin age, mute rate, runbook date. Fifteen minutes. Calendar, not vibes.
- Pin models = snapshot IDs in the node, not family aliases like
gpt-5.6oropus. Aliases move. Production should not.
What does “maintain” mean once the rail is live?
Maintain means the business outcome stays honest after the demo. The graph can stay green while the emails get worse, the CRM owner is wrong, and the model behind the node is a different snapshot than the one you evaluated. Maintenance is the loop that catches that without waiting for a customer to screenshot it.
Google SRE treats playbooks as part of on-call, not optional homework: when an alert fires, the human needs severity, impact, and the action that stops the bleeding (SRE Workbook, On-Call). Your n8n rail needs the same half-page, plus a pin that has not silently moved.
| Activity | Maintenance | Not maintenance |
|---|---|---|
| Owner | Named human can pause today | #ops with nobody on-call |
| Runbook | Date moved since the last vendor or prompt change | Launch doc in a private Drive |
| Errors | Shared handler with an execution link | Red row in Executions nobody opens |
| Review | Weekly, on a calendar, with a decision | “We’ll look when it breaks” |
| Models | Snapshot ID written on the runbook | Dropdown set to the family alias |
| Proof | Staging replay after a pin bump | Editor Execute on pinned fixture data |
Checklist before you call a live rail “handled”:
- Primary and backup still work here
- Runbook last-test date is newer than the last model, prompt, or vendor change
- Error workflow is attached on this graph, not only on the template you copied from
- Weekly review is on a calendar with a named chair
- Every AI node shows a snapshot ID, not
latestor a family alias
If any box is empty, you are renting uptime from luck.
Who owns a running automation on a normal Tuesday?
Ownership after go-live is the same job as ownership on exit week, minus the suitcase. The person who gets the alert must be able to pause, open the last failed execution, and either fix or escalate without asking Slack where the builder went. The complementary post covers handoff when that person leaves. Here the failure is they are still employed and nobody scheduled the work.
n8n project roles make the Tuesday test concrete. A project Admin can manage members, workflows, and credentials. An Editor can change the graph. A Viewer can look and cannot execute (n8n project roles). If your “owner” is a Viewer, they cannot pause. Fix the role before you need it.
| Role | Tuesday duty | Disqualifier |
|---|---|---|
| Primary owner | Triages the alert, pauses if blast radius is unclear, files the ticket | “I’ll ask Jamie” |
| Backup | Same within one business day, without the primary on Slack | The #ops channel |
| Exec sponsor | Funds a rebuild when the rail is haunted | Optional; never the only name |
| Builder | May still edit, under the owner’s change rule | Not the alert destination |
Tuesday owner checklist:
- Alert route is a shared inbox or on-call, not a personal Gmail
- Owner has Editor or Admin on the company project
- Backup has clicked Pause once in the last 90 days
- Change rule is written: staging first on anything that emails, charges, or writes a CRM
- New graphs inherit the same owner field, or they do not ship
Assumed access is not access. Across 500+ live rails, the ones that survive a vacation are the ones where a second person has already clicked Pause.
What stays on the runbook after go-live?
The launch runbook is a snapshot. Maintenance is keeping that snapshot true. Sticky notes on the canvas evaporate when someone rebuilds a node. Put the living copy in the workflow description plus a linked doc the backup can open at 2am.
Paste this block and update the dated rows when anything in it moves:
Name: [workflow id]
Owner / backup: [name] / [name]
Purpose: [one sentence]
Trigger: [webhook | cron | app event]
Irreversible steps: [emails, CRM writes, charges, deletes]
Pause procedure: [click path]
Credentials: [names — not values]
Model pin: [exact ID] · last eval: [date]
Prompt version: [id or hash]
Error workflow: [handler name] · last prove: [date]
Last staging test: [date + what was proven]
| Field that goes stale | How you notice | What you do |
|---|---|---|
| Model pin | Tone, JSON shape, or cost jumps with no prompt edit | Diff the node ID against the runbook |
| Prompt version | Someone “just tightened the system message” on prod | Revert, then change in staging with a date |
| Error workflow | New graph copied without Settings | Attach the shared handler before activate |
| Last test | Vendor renamed a field | Replay staging against the live schema |
| Owner | HR moved them last sprint | Reassign before the next alert |
Runbook hygiene checklist:
- Backup can find this without asking the primary
- Model pin matches the live node, character for character
- Irreversible steps still list every customer-facing write
- Last-test date moved after the last pin bump
- Pause procedure still works for someone who did not build it
If a row is blank, you do not understand the rail yet. Do not add a second feature on top of a blank irreversible-steps line.
How do I keep the error workflow paging next month?
n8n’s error-workflow docs are blunt: set an error workflow in Workflow Settings, and that workflow must start with an Error Trigger. It runs when an execution fails. You can force a failure with Stop And Error when a schema check should fail the run on purpose.
Facts from those pages that change maintenance, not just setup:
- You cannot test the Error Trigger by clicking Execute Workflow. It only fires on an automatic failure.
- If a workflow contains an Error Trigger, it uses itself as its error workflow by default.
execution.idandexecution.urlare missing when the trigger node itself failed.- Error-workflow executions do not count toward the execution quota. Do not skip the handler to “save runs.”
- New graphs do not inherit the handler. Copy-paste of nodes does not copy Settings.
The implementable alert contract lives in the handbook’s error-workflow spoke. For maintenance, the job is: still attached, still proven, still unmuted.
| Maintenance check | Pass | Fail |
|---|---|---|
| Attachment | Every active prod workflow names the shared handler in Settings | Template had it; this copy does not |
| Prove date | Automatic Stop And Error in staging this month | “It worked at launch” |
| Destination | Shared ops channel or on-call | Builder’s DM |
| Continue on Fail | Only on named enrichment branches | On the write node, hiding the failure |
| Mute rate | Alerts still get a human response | Channel is on mute, handler still firing |
Prove it on a cadence, given manual Execute will not fire the trigger:
- Keep a tiny staging workflow that hits Stop And Error on a webhook or schedule.
- Point its Error workflow at the shared handler.
- Trigger it via the automatic path, not the editor play button.
- Confirm the alert has workflow name, failed node, owner, and execution link (or an explicit missing-link line).
- Confirm a trigger-node failure still notifies when
execution.urlis absent. - Write the prove date on the runbook.
If you cannot name the handler in Settings on a graph you activated last week, you do not have error handling. You have hope plus a channel.
What does the weekly review actually inspect?
Keep it boring. Google SRE’s toil definition is the filter: work that is manual, repetitive, interrupt-driven, and scales linearly with traffic is toil, and they cap it so engineering still happens (Eliminating Toil). The weekly review is supposed to reduce next week’s toil — update a pin, kill an orphan, fix a mute — not become a second full-time job of clicking red rows.
n8n Insights gives owners a 7-day banner: production executions, failed production executions, failure rate, time saved (if you set it), run time average. It counts production runs (schedule, webhook), including error-workflow executions. It does not count manual editor runs or sub-workflow executions. Error-workflow outcomes are scored against the handler, not the parent that failed. Time saved is a number you typed in Settings, not a measured stopwatch.
Use Insights as the opening slide. Do not let it be the whole meeting.
| Weekly check | Pass | Fail |
|---|---|---|
| Failures | Named owner has a ticket or a mute reason | Unread error mail in a personal inbox |
| Zero runs | Expected silence, or a trigger bug you can name | “It should have fired” and nobody looked |
| Pin age | Snapshot ID matches runbook; deprecation not inside 60 days | Family alias, or a retired ID still in the node |
| Mute rate | Ops channel still gets a human reply | Sales muted the bot; handler still posting |
| Runbook date | Moved if anything changed | Launch date from the demo week |
| Orphan | One dated decision if you found a Copy of | List only grows |
Fifteen-minute procedure:
- Open Insights (7-day) plus Executions filtered to Error on the production project.
- Scan overnight failures and “zero runs” on rails that should have fired.
- Confirm the backup owner still works here.
- Diff model IDs in AI nodes against the runbook.
- Kill or pause one orphan if you find one.
- If a vendor mailed a deprecation, put a pin-bump on this week’s staging queue.
Sample agenda — same order every week, so a tired human can run it:
| Minute | Object | Decision you write down |
|---|---|---|
| 0–3 | Insights 7-day vs last week | Is failure rate a real incident or a noisy optional node? |
| 3–7 | Executions = Error, red rails first | Ticket, mute-with-reason, or pause |
| 7–10 | Zero-run list | Expected quiet, dead webhook, or unpublished by autodeactivation |
| 10–13 | Pin sheet vs live nodes | Alias found → staging bump this week, not “later” |
| 13–15 | One orphan or one runbook date | Dated kill/keep/reassign |
Sample weekly sheet (copy this; fill it in the meeting, not after):
| Workflow | Active | Last run | Failures 7d | Model pin | Runbook date | Decision |
|---|---|---|---|---|---|---|
revops-hubspot-lead-route-prod | yes | today | 2 | gpt-5.6-terra | 2026-07-06 | ticket #1842 |
support-draft-reply-prod | yes | today | 0 | claude-sonnet-5 | 2026-06-02 | pin 61 days old — eval this week |
Copy of invoice v2 | yes | 12d | 0 | — | — | pause 7d |
Decisions without dates are not decisions. A quiet Copy of that still sends is a maintenance miss even when Insights looks calm.
Self-hosted pruning defaults matter here. n8n deletes finished executions after EXECUTIONS_DATA_MAX_AGE hours — 336 hours, 14 days by default — or when the count exceeds EXECUTIONS_DATA_PRUNE_MAX_COUNT (10,000), oldest first (manage execution data). Annotated executions are kept. If your only forensic store is Executions, a monthly incident review is already gone. Export the DLQ and the alert log to something that outlives 14 days.
A lead routing rail is a good weekly teacher. If sales muted the channel, the graph can still be “green” while owners are wrong. Check mute rate and “why me?” CRM fields, not only failure rate.
Why pin models instead of leaving “latest” in the node?
Because the prompt file is not a freeze. Quality moves when the model string moves. OpenAI’s own model pages are explicit: snapshots let you lock a specific version so performance and behavior remain consistent (GPT-5.6 Sol). The family alias gpt-5.6 routes to Sol today (model guidance). That routing is their decision, not yours. Production should call gpt-5.6-sol or a dated snapshot, not the alias, unless you want the vendor to move you.
Anthropic’s lifecycle is the same idea with a clock. They notify customers with active deployments and give at least 60 days’ notice before retiring a publicly released model. They also tell you to export usage from the Claude Console and find leftover IDs before the retirement date (model deprecations). Family aliases such as opus and sonnet resolve to whatever the provider currently recommends; pin the full ID, for example claude-opus-5 or claude-sonnet-5 (model configuration).
| Identifier you typed | What it is | Safe in production? |
|---|---|---|
gpt-5.6 | Alias that currently routes to Sol | No — vendor can retarget it |
gpt-5.6-sol / dated snapshot | Explicit pin | Yes, until you bump it |
gpt-5.6-terra / gpt-5.6-luna | Explicit cheaper pins | Yes, if that is the evaluated tier |
opus / sonnet | Family alias | No |
claude-opus-5 / claude-sonnet-5 | Full ID | Yes, until you bump it |
| empty / first dropdown row | Whatever the credential loaded | No |
Failure mode we see on live AI rails: the node is set to a family alias. The vendor ships a new snapshot. Your JSON extractor starts wrapping keys differently. Nobody edited the prompt. Support tickets rise. Someone “fixes” it by adding a second instruction on production. Now you have drift plus an unversioned prompt.
Do not rewrite the prompt to paper over a moved pin. Score a small frozen set against the last known-good ID first. Then bump on purpose.
How do I pin models in n8n without freezing forever?
Pinning is not a religion against upgrades. It is a change-control rule: production does not float; staging may. n8n’s OpenAI Chat Model node dynamically loads models from OpenAI for the credential you attached. The first row in that dropdown is not a pin. Type or select the snapshot ID. Same rule on Anthropic, Google, and HTTP Request bodies that send "model": "...".
| Place the string lives | How you pin | How it drifts |
|---|---|---|
| OpenAI / Anthropic Chat Model sub-node | Set Model to the snapshot ID | Re-open the node and pick the alias again |
| OpenAI app node (text operations) | Same Model field | Template you imported still has the alias |
| HTTP Request to the vendor API | JSON body "model" | Expression that reads $env.LATEST_MODEL |
| Code node | Constant at the top of the file | Hardcoded alias in a copied snippet |
| Agent / chain root | The attached language-model sub-node | You pinned the root and forgot the sub-node |
Pin-and-bump procedure:
- Inventory every AI node: workflow, node name, current model string, credential.
- Replace family aliases with snapshot IDs. Write the ID on the runbook.
- Log
model_idon every execution (Set node into your ops table). If the log and the node disagree, the node is not what you think. - Keep a 20–50 example golden slice for that rail. Replay it in staging on the current pin.
- When you want a new snapshot: replay the slice on the candidate in staging. Compare JSON shape, refusal rate, and cost. Then bump production and move the runbook date.
- Watch vendor deprecation mail. Anthropic’s 60-day floor is a calendar item, not a vibe. OpenAI publishes a deprecations page — check it in the same weekly review, not when the node starts 404ing.
- No production AI node uses
latest, a family alias, or a blank Model field - Runbook pin matches the live node
- Golden slice exists for any rail that emails or writes a CRM
- Pin bumps go through staging, not a Friday dropdown click
- Usage export (OpenAI dashboard / Claude Console) has been grepped for retired IDs this quarter
A pin with no bump process is how you ride a snapshot into retirement and learn about it from a hard API error on a customer send.
What else drifts besides the model?
The model string is the loud one. The quiet ones take the same weekly pass.
| Drift | Symptom | Maintenance move |
|---|---|---|
| Prompt text | Tone or policy changes with the same pin | Version the system prompt; forbid prod edits |
| Tool / JSON schema | Model calls a renamed field; write node nulls it | Schema check before the irreversible node |
| Vendor API | HubSpot / Stripe field or scope change | Staging replay after the changelog |
| Credentials | 401 after a password reset or app reinstall | Company-owned credential; rotate on a calendar |
| Continue on Fail | Parent looks successful; enrichment is empty | Error workflow never fires; check the skip log |
| Traffic mix | Eval still green; live tickets got weirder | Sample live traces, not only the golden slice |
| n8n version | Node defaults changed on instance upgrade | Read the changelog before you click upgrade |
| Insights time saved | Banner looks like ROI | You typed the minutes; they are not a measurement |
Prompt versioning checklist:
- System prompt lives in a Code node constant, a config table, or source control — not a sticky note
-
prompt_versionis logged next tomodel_id - Production edits to the prompt require the same staging replay as a pin bump
- “Quick wording fix” on the live Basic LLM Chain is a policy violation, not a flex
Credentials are a maintenance item even when nobody is leaving. OAuth still dies when an admin revokes a connected app, a scope is removed, or a refresh token expires. The ownership post is the cutover playbook. The weekly review is “did any 401 cluster appear overnight?”
Self-hosted operators should also know N8N_WORKFLOW_AUTODEACTIVATION_ENABLED exists and defaults to false. If someone turns it on, n8n can unpublish a workflow after repeated crashed executions (N8N_WORKFLOW_AUTODEACTIVATION_MAX_LAST_EXECUTIONS, default 3) (executions env). That is a silent pause. Put “is it still published?” on the zero-run check.
How do I implement this maintenance loop in n8n?
Do not build a second product called “the maintenance bot” on day one. Wire the five controls onto the rails you already run.
| Control | n8n object | Done when |
|---|---|---|
| Owner | Project role + workflow description | Backup can pause without a screen share |
| Runbook | Description + linked doc | Dated rows match the live graph |
| Error workflow | One Error Trigger workflow, attached in Settings | Automatic prove this month |
| Weekly review | Calendar event + Insights + Executions | 15 minutes, dated decisions |
| Model pin | Snapshot ID on every AI node + log field | Inventory sheet has no aliases |
Stand-up sequence for one production project:
- Export the active workflow list (name, active, last run, last editor).
- Attach the shared error handler under Settings on every active graph. New graphs get a definition-of-done checkbox for this.
- Fill the runbook block, including model pin and prompt version, on every rail that emails, charges, or writes a CRM.
- Add a Set node (or the first Code node) that writes
model_id,prompt_version,workflow_id, andexecution_idto your ops table. - Create the 15-minute weekly calendar. Same weekday. Named chair. Not “when we have time.”
- Inventory AI nodes. Replace aliases. Replay one golden slice in staging.
- Decide Insights time saved only if you want a banner. Do not use it as the health metric.
Inventory the AI nodes the same week you attach the handler. One row per Model field, including sub-nodes hiding under an Agent or chain:
| Workflow | Node | Current string | Pin? | Last eval | Action |
|---|---|---|---|---|---|
support-draft-reply-prod | OpenAI Chat Model | gpt-5.6 | no — alias | never | set gpt-5.6-terra, replay 25 gold rows |
support-draft-reply-prod | HTTP classify | claude-sonnet-5 | yes | 2026-07-01 | keep |
revops-lead-enrich-prod | Anthropic Chat Model | sonnet | no — alias | never | set claude-sonnet-5, stage first |
ops-digest-internal | OpenAI Chat Model | gpt-5.6-luna | yes | 2026-06-20 | keep; cheaper tier is the point |
If a row has no last eval, it is not pinned. It is hoped. Do not activate a new customer-email path until that sheet has a date on the pin.
| Do this in n8n | Do not do this |
|---|---|
| One shared Error Trigger handler | A Slack node at the end of each graph |
| Snapshot ID typed in Model | First dropdown row after reconnecting the credential |
| Staging webhook prove of the handler | Editor Execute as the prove |
| Ops table outliving 14-day prune | Executions UI as the only archive |
| Source control / second project for edits | Live-edit a customer email path on Friday |
Paid plans add Git-backed environments. That still does not pin models for you. The string in the node is the pin. Environments just give you a less stupid place to bump it.
If the rail is lead routing, add mute rate and “owner reason written on the CRM row” to the weekly checks. A silent wrong owner is a maintenance miss even when Insights shows a 0% failure rate.
What usually fails first when teams skip maintenance?
The first break is almost never “the canvas crashed.” It is a quiet wrong send.
Typical stack, in order:
- No owner on the alert. Handler posts to the builder’s DM. Builder is in a meeting. Customer gets the duplicate.
- Family alias in the Model field. Vendor moves the alias. Extractor JSON shifts. Downstream CRM write stores empty fields. Graph still green.
- Continue on Fail on the write-adjacent node. Error workflow never fires. Insights failure rate stays pretty.
- New workflow copied without Settings. No error workflow. Failures exist only in Executions.
- Weekly review skipped for three weeks. Pin is 90 days old. Deprecation mail sat in a personal inbox. Node starts erroring on the retirement date.
- Executions pruned at 14 days. You cannot reconstruct what happened last month. The runbook still says “last test: launch.”
| Skip | What breaks | Cost shape | Instead |
|---|---|---|---|
| Named owner | Alerts die in a DM | Hours of “who is on this?” | Two humans, Editor role |
| Current runbook | Backup guesses irreversible steps | Double email, double charge | Dated half-page |
| Error workflow attach | Failures are diary entries | Customer reports first | Shared handler + monthly prove |
| Weekly review | Drift compounds | A week of junk before anyone looks | 15 minutes on a calendar |
| Model pin | Quality moves with no prompt edit | Support load, then a panic prompt edit | Snapshot ID + staging replay |
Concrete incident shape: an AI drafting node on a customer-email rail sits on gpt-5.6. The alias retargets. The model starts adding a second CTA the extractor does not expect. The Send node still fires. Nobody changed a prompt. Sales mutes the thread. You find it in the weekly review only if you look at mute rate and sample outputs, not if you only look at failure rate.
Continue on Fail is the accomplice. n8n treats the execution as successful if you swallowed the error, so the Error Trigger does not run. That is documented community behavior and it matches the product: the handler fires on a failed execution, not on a branch you marked as fine. Use Continue on Fail for optional enrichment. Do not use it on the node that talks to customers.
How do I measure whether maintenance is working?
Measure the loop, not the canvas. I will not invent a studio-wide MTTR or a “% of workflows healthy” number for your instance. Those only exist after you count your tickets, your mute rate, and your pin ages.
| Signal | How you count it | Healthy | Lying |
|---|---|---|---|
| Time to pause | Stopwatch on a staging drill, twice a year | Backup pauses without a screen share | “They could if they tried” |
| Alert response | First human reply on the ops thread | Minutes to hours, with a named owner | Unread count climbing |
| Pin inventory | Spreadsheet vs live nodes | Zero family aliases in prod | Dropdown “looks current” |
| Runbook age | Date vs last vendor/prompt/model change | Date moved when the rail moved | Launch date forever |
| Handler prove | Automatic Stop And Error this month | Alert fields still complete | “We tested it in the editor” |
| Silent zero-runs | Expected volume vs actual | You can explain a quiet day | Insights 7-day average hides a dead webhook |
| Mute rate | Sales/ops still reading the bot | Thread gets replies | Channel muted, graph green |
| DLQ age | Oldest unreplayed item | Bounded, with a weekly burn-down | Queue nobody opens |
Insights is allowed as a dashboard, with caveats already named: 7-day default, production executions only, error-workflow stats land on the handler, time saved is configured not measured (Insights). Pair it with the ops table you write from the Set node. If you do not log model_id, you cannot prove the pin next month.
Quarterly extras:
- Export Claude / OpenAI usage and grep for IDs you thought you retired (Anthropic usage export).
- Re-prove the Error Trigger on an automatic path.
- Re-open the orphan sheet. Dated decisions, not “later.”
- Confirm pruning and Insights retention still match how long you need history.
Use this in the owner 1:1, not as wallpaper.
| Question | Pass |
|---|---|
| Who is primary / backup? | Two living humans, Editor or Admin |
| Can backup pause it today? | Demonstrated this quarter |
| Error workflow named in Settings? | Shared handler, prove date < 30 days |
| Model string on every AI node? | Snapshot ID, matches runbook, logged on the run |
| Runbook last updated? | Since the last pin, prompt, or vendor change |
| Weekly review happening? | Calendar + last sheet dated |
| Mute rate known? | Ops/sales still reading, or a written mute reason |
| DLQ / ops table outlives prune? | Older than 14 days still searchable |
Fail any row on a customer-facing rail → fix before adding scope. A new extractor on an unpinned, unpaged graph is how you multiply the blast radius.
If pin inventory, handler prove, and mute rate are all unknown, maintenance is not working. You have a blog post in a folder.
When should I hire vs DIY this?
DIY the loop when the writes reverse in an hour, the owner already sits in the company, and you can pause without a customer seeing it. Hire (or book the audit) when the rail talks to customers or money and nobody on staff will own the Tuesday review.
| Situation | DIY | Hire / audit |
|---|---|---|
| Internal digest, reversible | You can attach the handler and pin models this week | You have not opened n8n since the agency left |
| Lead routing, CRM write | You have a named sales ops owner and a staging project | Wrong owner already reached a customer |
| Customer email drafted by a model | Only if pins, golden slice, and approval gate exist | Alias in the node, no sample review |
| Charges / refunds | Almost never DIY as the first live AI rail | Spine first: idempotency, DLQ, owner, pin |
| 40 undocumented Zaps plus one n8n AI graph | You will spend the week on inventory, not features | You need a kill/keep pass before more canvas |
DIY week, if that is the honest call:
- Named primary and backup with Editor access
- Shared error workflow attached and proven on an automatic path
- Model pins inventoried
- Weekly 15-minute slot created
- One golden slice for the scariest rail
Book the $500 Automation Audit when you cannot name the owner, the pin, or the pause path on a rail that already emails customers. Bring the workflow list, not a demo recording.
What should I skip if I only have a week?
Skip new agents, queue mode, and a custom observability stack. Do the five controls on the rails that can hurt people.
| This week | Not this week |
|---|---|
| Name owner + backup; fix roles | A new AI feature on an unowned rail |
| Attach the shared error workflow; prove it once | Per-graph Slack copy |
| Inventory model strings; kill aliases on customer-facing rails | Fine-tuning, prompt playgrounds, extra models |
| Half-page runbook on money / email / CRM writes | A 12-page Confluence theme |
| Calendar the weekly 15 minutes | A full Insights rollout as the “program” |
Pause undocumented Copy of graphs that send | Rebuilding the haunted god-workflow |
One-week procedure:
- Export active workflows. Mark red: customer email, CRM write, money.
- Put two human names on every red row. If you cannot, pause the row.
- Attach the shared Error Trigger handler. Prove it with Stop And Error on a staging webhook.
- Open every AI node on red rows. Replace aliases with snapshot IDs. Write them down.
- Fill the runbook block on those rows. Blank irreversible-steps = pause.
- Create the weekly calendar event. Put Insights + Executions + the pin sheet in the invite.
A week of that beats a week of canvas. Features on an unowned, unpinned, unpaged rail multiply blast radius. Maintenance is how you earn the right to add the next node.
FAQ
How do I manage and maintain AI automations once they’re running?
Name an owner who can pause the workflow, keep a half-page runbook current, attach a shared n8n error workflow, run a weekly review, and pin snapshot model IDs in every AI node. A green canvas is not the program. The complementary job — what happens when the builder leaves — is ownership and credential cutover, not this Tuesday loop.
How do I measure whether maintenance is working?
Count pin inventory (zero family aliases in production), error-workflow prove date, mute rate, silent zero-runs you can explain, and runbook age versus the last real change. Insights failure rate is a start and it lies if you Continue on Fail, if you only look at the handler’s stats, or if time saved is a number you typed. If you cannot name those signals, the loop is not running.
What usually fails first when teams try this?
The alert destination and the model string. Handlers still post to a personal DM, and AI nodes still sit on a family alias, so the first incident is a quiet wrong send rather than a crashed canvas. Continue on Fail on a write-adjacent node is the usual accomplice: the execution looks successful, so the Error Trigger never fires.
How long does this take to show results?
The controls themselves take days, not quarters: attach the handler, pin the models, name the owner, calendar the review. You feel it the first time a staging Stop And Error pages the right human, or the first weekly pass that catches a zero-run before a customer does. I will not invent a payback week for your book of work.
What should I skip if I only have a week?
Skip new agents, queue mode, and a custom dashboard. Spend the week on red rails only: owners, shared error workflow plus one automatic prove, snapshot pins on customer-facing AI nodes, half-page runbooks, and a recurring 15-minute review. Pause undocumented senders instead of rebuilding them.
When is this not worth doing yet?
When nothing is live, or the only graphs are reversible internal experiments with no customer send. You still want an owner and a pin before you activate a customer path. If you cannot name a human who will run the weekly review next quarter, do not add another AI node. Fix ownership first.
CTA
Green is not maintained. Name the owner, pin the model, page the human.
For the rest of the production spine, keep the handbook open, then use automation or book the audit.
What questions does this article answer?
- How do I manage and maintain AI automations once they're running?
- Name an owner who can pause the workflow, keep a half-page runbook current, attach a shared n8n error workflow, run a weekly review, and pin snapshot model IDs in every AI node. A green canvas is not the program. The complementary job — what happens when the builder leaves — is ownership and credential cutover, not this Tuesday loop.
- How do I measure whether maintenance is working?
- Count pin inventory (zero family aliases in production), error-workflow prove date, mute rate, silent zero-runs you can explain, and runbook age versus the last real change. Insights failure rate is a start and it lies if you Continue on Fail, if you only look at the handler’s stats, or if time saved is a number you typed. If you cannot name those signals, the loop is not running.
- What usually fails first when teams try this?
- The alert destination and the model string. Handlers still post to a personal DM, and AI nodes still sit on a family alias, so the first incident is a quiet wrong send rather than a crashed canvas. Continue on Fail on a write-adjacent node is the usual accomplice: the execution looks successful, so the Error Trigger never fires.
- How long does this take to show results?
- The controls themselves take days, not quarters: attach the handler, pin the models, name the owner, calendar the review. You feel it the first time a staging Stop And Error pages the right human, or the first weekly pass that catches a zero-run before a customer does. I will not invent a payback week for your book of work.
- What should I skip if I only have a week?
- Skip new agents, queue mode, and a custom dashboard. Spend the week on red rails only: owners, shared error workflow plus one automatic prove, snapshot pins on customer-facing AI nodes, half-page runbooks, and a recurring 15-minute review. Pause undocumented senders instead of rebuilding them.
- When is this not worth doing yet?
- When nothing is live, or the only graphs are reversible internal experiments with no customer send. You still want an owner and a pin before you activate a customer path. If you cannot name a human who will run the weekly review next quarter, do not add another AI node. Fix ownership first.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation Why doesn’t worker concurrency cap my n8n sub-workflows
Worker concurrency does not cap n8n sub-workflows. Each Execute Workflow child is a new execution the production limit skips, usually on the parent worker.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.