Why doesn’t worker concurrency cap my n8n sub-workflows
Worker concurrency does not cap n8n sub-workflows. Each Execute Workflow child is a new execution the production limit skips, usually on the parent worker.
William Spurlock Founder — Spurlock Studios 34 MIN
Worker concurrency does not cap n8n sub-workflows because the production limit only counts webhook and trigger starts. Each Execute Sub-workflow call still opens a new execution — you get a child ID and a View sub-execution link — but that child is not a production start. Official concurrency control lists sub-workflow executions as exempt, same as manual, error, and CLI runs.
In queue mode the miss is worse: as of August 2026, those children typically keep running on the worker that picked the parent. They are usually not new Redis jobs, so --concurrency never sees them. Parent slot ≠ child work. The Production n8n handbook is the spine. This spoke is the cap you thought you set.
The short answer
- The production cap counts trigger starts, not child executions.
N8N_CONCURRENCY_PRODUCTION_LIMITand Cloud plan concurrency skip sub-workflows by design. - A new execution ID is not a new queued job. Execute Sub-workflow records a child run. In queue mode it typically stays inline on the parent worker.
- Worker
--concurrencylimits Bull jobs per worker. Default is 10. A parent that occupies one job can still run dozens of children on that process. - Wait for Sub-Workflow Completion does not enroll children in the cap. It only decides whether the parent holds its slot until the child finishes.
- To actually bound child work, batch items, wait inside the graph, use one consumer, or start children via webhook so they become production jobs. Then put idempotency keys on those children.
Why doesn’t worker concurrency cap n8n sub-workflows?
Because n8n’s limiter is keyed off how the run started, not off how many execution rows you see. A Schedule or Webhook start is a production execution. An Execute Sub-workflow start is a different kind. The controller is documented to ignore it.
That is easy to miss in the UI. The child has its own execution. It has its own status. It can fail independently. Operators reasonably assume “executions” and “concurrency slots” are the same pile. They are not.
| Start path | New execution row? | Counts toward production concurrency? | Typical queue-mode job? |
|---|---|---|---|
| Webhook / Schedule / app trigger | Yes | Yes (if the controller is on) | Yes — main enqueues, a worker picks it |
| Editor Execute Workflow | Yes | No — manual | Usually main, not a worker |
| Execute Sub-workflow / Execute Workflow node | Yes — child ID | No — exempt | Usually no — inline on parent worker |
| Error workflow | Yes | No — exempt | Outside the production cap |
| CLI start | Yes | No — exempt | Depends on how you invoked it |
Procedure to see the exemption on your box, not in a blog:
- Set
N8N_CONCURRENCY_PRODUCTION_LIMIT=1(or--concurrency=1in queue mode). Pick one source of truth. If the env is not-1, n8n takes the worker limit from the env and ignores a conflicting flag. - Publish a parent that fans Run once for each item into 20 Execute Sub-workflow calls. Put a 10-second Wait in the child.
- Fire the parent from a published webhook or schedule. Do not click Execute in the editor.
- Open Executions. You should see one parent plus many children overlapping in time.
- If 20 children run while the cap still says 1, the limiter is doing what the docs say. It is not broken. You aimed it at the wrong object.
If step 3 was a manual click, you proved nothing. Manual is also exempt.
View sub-execution proves a child run exists. It does not prove the child was a queued job. n8n documents the link both ways: parent node → child execution, child execution → parent (Execute Sub-workflow). Use the link for debugging. Do not use it as evidence the worker pool load-balanced anything.
| What you clicked | What it proves | What it does not prove |
|---|---|---|
| View sub-execution | A child execution ID was created | That Redis got a second job |
| Child status success/error | That graph finished | That it counted against the cap |
| Parent still running while children appear | Wait is on, or children are slow | That workers split the tree |
| Settings → Workers (if you have it) | Which process is hot | Nothing, if the UI is not on your plan |
Settings → Workers is documented as self-hosted Enterprise, and on n8n Cloud Enterprise you contact n8n to enable it (queue mode). No Workers screen does not mean you have no pin. It means you watch host CPU per container instead.
What does N8N_CONCURRENCY_PRODUCTION_LIMIT actually count?
It counts concurrent production executions started from a webhook or trigger. Default is -1 (off) in regular mode. Set a number and excess production starts wait FIFO. You cannot retry a queued execution; cancel or delete drops it from that wait list. On startup, n8n resumes up to the limit and re-enqueues the rest (control concurrency).
Queue mode uses a separate mechanism — how many jobs one worker may run — but the same env var, if it is not -1, overrides --concurrency. Worker --concurrency defaults to 10. n8n recommends 5 or higher; very low concurrency with many workers can exhaust the database connection pool (queue mode).
| Lever | What it limits | What it does not limit |
|---|---|---|
N8N_CONCURRENCY_PRODUCTION_LIMIT in regular mode | Concurrent webhook/trigger starts on the instance | Sub-workflows, manual, error, CLI, HTTP calls inside a run |
--concurrency on n8n worker | Bull jobs that worker pulls | Inline children, other workers, vendor req/s |
Same env var in queue mode (not -1) | Overrides the worker flag | Still not sub-workflow executions |
| Cloud plan concurrency | Production starts on that Cloud instance | Sub-workflows and error runs; exact plan numbers live on pricing |
| Execution quota | Billable production runs | Sub-workflow children — n8n counts only the parent toward quota (executions) |
Quota and concurrency are different meters. Both skip children. Do not treat “this child did not increment my plan usage” as proof the child was cheap on CPU.
Checklist before you trust a number you already set:
- You know whether the instance is regular mode or
EXECUTIONS_MODE=queue - You have one source of truth: env or
--concurrency, not both fighting - The cap was tested with a published trigger, not a canvas click
- You looked at child overlap, not only parent queue depth
- You did not confuse quota remaining with slots remaining
A cap that only serializes parents is a parent cap. Name it that in the runbook.
Parent concurrency is not child executions
Parent concurrency is the number of top-level production jobs allowed to run. Child executions are the rows Execute Sub-workflow creates under those jobs. One parent can own many children. The cap looks at the parent. The vendor looks at the children.
This is the sentence to put on the whiteboard: one occupied slot can still mean N overlapping child graphs.
Worked shape, not a benchmark:
| Graph | Parent slots in use | Child execution rows | HTTP calls the vendor sees |
|---|---|---|---|
| One webhook, no sub-workflow, five sequential HTTP nodes | 1 | 0 | 5, one after another |
| One webhook, Execute Sub-workflow Run once with all items, child does five HTTP nodes | 1 | 1 | 5, on the child |
| One webhook, Run once for each item, 50 items, child does one HTTP node | 1 | 50 | 50, overlapping unless you Wait |
| Five webhook parents, no children, cap=5 | 5 | 0 | Whatever those five graphs do |
The third row is where teams get surprised. They set workers to 5 because Stripe or Airtable “can take 5.” The parent is 1. The children are 50. The limiter never saw 50.
Decision list when you design the split:
- Is the child a mapping helper (clean fields, shared Code)? Keep Execute Sub-workflow. Bound items in the parent.
- Is the child a write (CRM, mail, charge, delete)? Bound items and put the idempotency claim in the child, not as a comment on the parent.
- Is the child a fan-out you want distributed across workers? Execute Sub-workflow will not do that in current queue-mode architecture. Use a production trigger on the child.
- If you cannot pick 1–3 in one sitting, do not split yet. A fat parent with a Wait is safer than an unbounded tree.
The split is for reuse and testing. It is not a concurrency primitive.
Also check This workflow can be called by on the child (Workflow settings). A helper that anyone can call from a second parent doubles the fan-out you just bounded on the first parent. Two published callers, one unbounded child, one vendor. The child’s settings are part of the cap.
| Child setting | Healthy | Dangerous |
|---|---|---|
| Called by one named parent | You can count callers | “Anyone” plus two more parents next quarter |
| Input data mode defined | Schema is a contract | Accept all data into a write |
| Save successful production executions | You can count child rows | You cannot prove overlap after prune |
If you cannot list the callers, you cannot size the child.
Queue mode: children stay on the parent worker
Queue mode is: main (or a webhook processor) accepts the trigger, Redis holds the job, a worker pulls it, the worker loads the workflow from Postgres, the worker runs it (enable queue mode). That story is about the parent job.
Execute Sub-workflow, as of August 2026, typically does not put a second job on Redis. n8n staff have said the same on the community: if you start children with the sub-workflow node, the chain stays on the worker that picked the parent; webhook-triggered children can distribute (forum thread). Treat that as current architecture, not a forever guarantee. Confirm on your version with the measurement section below.
Hedge, because this is the part that moves: n8n has renamed the node (Execute Workflow → Execute Sub-workflow), shipped 2.x queue internals, and changed how stop signals reach children. A later build could enqueue children. Until Redis job count rises 1:1 with child execution IDs on a fan-out, assume inline.
| Pattern | Distributes across workers? | Cap that applies |
|---|---|---|
| Parent + Execute Sub-workflow (wait on or off) | No — same worker, typically | Parent job only |
| Parent HTTP → child production webhook | Yes — each child is a job | Production / worker concurrency on those jobs |
| Nested sub-sub-sub tree | No — still the first worker | None on the descendants |
| Error workflow on any of the above | Outside the production cap | Do not use it as throughput |
Failure picture: five workers, queue depth looks fine, one worker pegged, four idle. You buy a sixth worker. The hot worker stays hot because the tree never left it.
Do not debug that with average CPU across the pool. Average hides the pin.
Nested depth makes the pin longer, not wider. Parent → child → grandchild via Execute Sub-workflow is still one worker’s job from the queue’s point of view. A one-minute tree is one minute on that process. Adding a fourth worker does not take the grandchild.
AI Agent tool calls that invoke a workflow sit in the same family in n8n’s own queue-mode bug reports: the tool run is a sub-workflow-style child, not a fresh Bull job. Hedge: agent tool wiring moves faster than core Execute Sub-workflow. If a tool-call child shows a separate job ID in Redis on your build, believe Redis. If it does not, treat it as inline.
Stopping is version-specific. Older queue-mode builds could mark a child canceled in the database while the worker kept running it, then overwrite the status back to success — because there was no job for the abort path to target. Newer builds added stop propagation. Do not assume Stop in the UI halted child HTTP. Watch the vendor. Confirm on the version string in Settings.
- You recorded n8n version on main and every worker (they must match)
- You know whether children appear as Redis jobs on this build
- Nested Execute Sub-workflow depth on the hot path is 1, not 4
- Agent-tool children were checked the same way as Execute Sub-workflow children
- Stop was verified against the downstream API, not only the execution row
Mismatching worker versions is how you get “it distributes on staging” and “it pins in production” in the same week.
Wait for Sub-Workflow Completion is not a cap
The node option Wait for Sub-Workflow Completion decides whether the parent pauses until the child returns data. It does not enroll the child in N8N_CONCURRENCY_PRODUCTION_LIMIT. It does not create a Bull job. It does not rate-limit HTTP inside the child.
| Wait option | Parent behavior | Child behavior (typical, as of August 2026) | Slot picture |
|---|---|---|---|
| On (default in most canvases) | Parent holds until child finishes | Child runs inline on that worker | One job slot; child CPU on the same process |
| Off | Parent continues; may finish while children still run | Still typically same worker, not a queued sibling | Parent slot may free while child work continues on the process |
Off is not “queue the children.” Community reports of Wait=false plus --concurrency=1 still running many children are consistent with the exemption: the children were never jobs the worker flag could refuse.
Use Wait for data you need downstream. Use a Wait node or batch size for pacing. Do not mix those two ideas.
Procedure if you need the parent to collect child results and stay gentle on a vendor:
- Keep Wait for Sub-Workflow Completion on so you get return data.
- Do not set Mode to Run once for each item on an unbounded list.
- Batch upstream (Loop Over Items / Split In Batches on older graphs) to a size you measured.
- Put a Wait in the child or between batches in the parent.
- Confirm overlapping child timestamps in Executions still match the batch size, not the full list.
If you turn Wait off so the parent “stays fast,” you still owe the child a bound. Fast parents with unbounded children are how a “simple sync” knocks Airtable into a 30-second lockout. Queue mode made that easier, not safer.
Run once for each item multiplies executions
Execute Sub-workflow Mode is the hidden multiplier.
| Mode | Child executions per parent run | When it is honest |
|---|---|---|
| Run once with all items | One child, all items in | Mapping, reduce, shared lookup |
| Run once for each item | One child per item | You have already capped how many items exist |
“Per item” is not parallel-with-a-cap. It is “start a new exempt execution for every row.” Combine it with a 2,000-row Google Sheet and you have 2,000 child graphs on one worker process. The parent still counts as one production start.
Loop Over Items in the parent plus one child call per batch is a different shape: you control how many child executions exist. Loop Over Items plus per-item Execute Sub-workflow inside the loop is the same explosion with extra nodes.
Checklist before you ship per-item mode:
- Max items per parent run is named (pagination, query limit, or hard slice)
- Child HTTP is bounded (Wait, batching, or a single-consumer workflow)
- Child writes claim an idempotency key before the side effect
- You have watched one production parent and counted child rows
- Error workflow is attached on the child if the child can fail closed — see error workflows operators actually read
If you cannot name max items, you do not have a mode preference. You have a stampede setting.
Loop Over Items (Split In Batches on older exports) does not inherit worker concurrency either. It only slices the item list inside one execution. That is useful. It is still not --concurrency.
| Shape | Child executions | Parallelism you actually got |
|---|---|---|
| Per-item Execute Sub-workflow, 200 rows | 200 exempt children | Up to 200 overlapping graphs on one process |
| Loop 20 → one Execute Sub-workflow per loop, Wait between loops | 10 children, serial batches | 10 child runs, paced |
| Loop 20 → per-item Execute Sub-workflow inside the loop | Still ~200 children | Same stampede, extra loop node |
| One child, Run once with all items, HTTP batching inside the child | 1 child | Whatever that HTTP node does |
Exported JSON is the audit when the UI renamed the loop node. Search splitInBatches and the current Loop Over Items type. Confirm on a node you just added, because n8n can rename fields without a blog post.
Procedure to convert a per-item bomb without changing the child’s contract:
- Note the child’s expected input (one item vs a list). If it is built for one item, keep one item — just do not start 200 executions at once.
- Add Loop Over Items in the parent with a batch size you can defend (start at 5 or 10, not 200).
- Inside the loop: Execute Sub-workflow Run once with all items for that slice, or one item at a time after the previous child returns.
- Wait between loops if the vendor needs it.
- Count child rows on the next production parent. It should match loop iterations (or iterations × slice size if you still per-item inside a small slice), not the raw list length.
If the child must stay one-item-in, serialize. Parallel-and-hope is how you meet the CRM’s lockout.
Regular mode, queue mode, and n8n Cloud are different caps
People mash three systems into “we set concurrency to 5.”
Regular mode (self-hosted): one process. N8N_CONCURRENCY_PRODUCTION_LIMIT FIFO-queues extra production starts. Default off. Sub-workflows exempt.
Queue mode (self-hosted): Redis + Postgres + workers. --concurrency (or the env override) is jobs per worker. Sub-workflows typically inline. Queue mode is not supported on SQLite. Share N8N_ENCRYPTION_KEY across main and workers.
n8n Cloud: plan concurrency in regular mode; exact numbers are on pricing and can change. Same exemption: production starts only, not sub-workflows, manual, or error runs (Cloud concurrency). Queue mode on Cloud is Enterprise and you contact n8n to enable it. Do not copy a self-hosted Docker compose onto Cloud and expect --concurrency to exist.
| Hosting | Where you set the number | Applies to sub-workflows? | Notes |
|---|---|---|---|
| Self-hosted regular | N8N_CONCURRENCY_PRODUCTION_LIMIT | No | Off until you set it |
| Self-hosted queue | --concurrency and/or the same env if not -1 | No, typically | Default 10 per worker |
| n8n Cloud regular | Plan limit, visible on the Executions tab | No | Do not invent the plan number here |
| n8n Cloud queue | Enterprise, contact n8n | Confirm on that enablement | Do not assume self-hosted inline behavior until you measure |
Evaluation / test-run concurrency is a fourth limiter (N8N_CONCURRENCY_EVALUATION_LIMIT / per-plan evaluation parallelism). Do not tune production workers by watching an evaluation slider.
If you are on Cloud Starter/Pro without queue mode, “add a worker” is not an available move. Bound the graph. That is the whole fix.
Same encryption key, same Postgres, same Redis. Workers that cannot decrypt credentials will fail children in ways that look like “concurrency.” They are not. N8N_ENCRYPTION_KEY must match main. Queue mode on SQLite is not a posture; n8n tells you not to.
| Drift | Looks like | Actually |
|---|---|---|
| Worker missing the encryption key | Child HTTP credential errors under load | That worker cannot read credentials |
| Main on 2.x, worker on 1.x | “Only some children distribute” | Mixed binaries; undefined behavior |
N8N_CONCURRENCY_PRODUCTION_LIMIT=20 and --concurrency=5 | Arguments in Slack | Env wins if it is not -1 |
| Cloud UI limit vs a compose file you keep in git | “We set 5” | Cloud is not reading your compose |
Write the actual host and the actual limiter in the workflow description. “Concurrency 5” with no host is how the next person retunes the wrong knob.
What fails first: the hot-worker landmine
The first failure is rarely a red execution. The first failure is a vendor 429, a stuck worker, or a green parent whose children still hammer an API after you thought the cap held them.
Concrete mode, no invented rates:
You set --concurrency=5 on three workers because the CRM “allows five connections.” A nightly parent loads 400 deals and Execute Sub-workflow per deal. One worker draws the parent job. Four hundred child executions run on that process. The other two workers idle. The CRM sees hundreds of overlapping HTTP calls from one Node process. The production cap still reads as healthy: one parent job.
Second-place failure: you “fix” it by adding workers. Cost goes up. The pin does not move.
Third-place failure: you switch children to webhooks so they distribute, forget idempotency, and the at-least-once queue double-writes. That is a different outage. Keys belong in the child before the write.
| Symptom | Likely cause | Wrong fix | Right fix |
|---|---|---|---|
| One worker 100% CPU, siblings idle | Execute Sub-workflow tree pinned to the first worker | Add workers | Shallow the tree or webhook-fan-out with keys |
| Vendor 429 with “concurrency=5” | Children exempt; per-item mode | Raise the cap | Batch + Wait, or one consumer workflow |
| Queue depth low, latency high | Inline children not visible as jobs | Scale Redis | Measure per-worker CPU and child overlap |
| Parents queued, children never start | Webhook children + Wait + all slots held by waiting parents | Raise concurrency blindly | Leave slots for children; cap parent fan-out |
| Duplicate CRM rows after the “fix” | Webhook children retried | Turn Wait off | Idempotency keys before the write |
Deadlock is the webhook-child version, not the Execute Sub-workflow version. If every worker slot is a parent waiting on HTTP children, and those children need a slot to run, nobody moves. Execute Sub-workflow inline work cannot deadlock that way because the child does not need a second slot. It can still melt the process.
Do not page from queue depth alone. Page from vendor 429s, worker CPU, and child overlap.
A fourth failure: the parent is green because Wait is off, the child later 429s, and nobody is watching child executions. Error workflows on the parent do not fire for a child that fails after the parent already succeeded. Attach the handler on the child. The error workflow spoke is the attach and the payload. This spoke is why a parent cap will not throttle the child that pages.
| Parent Wait | Child later fails | Who gets the error workflow? |
|---|---|---|
| On | Child error can fail the parent | Parent handler, if attached and the run is automatic |
| Off | Parent already succeeded | Child handler only — parent stays green |
| On, parent Continue on Fail | Child error swallowed | Often nobody |
Continue on Fail on Execute Sub-workflow is a mute on the split. If the child wrote, you now have a green parent and a partial apply. Fail closed on writes.
How to cap child work on purpose
You cap children with graph shape, not with a worker flag. Pick one pattern and write it in the workflow description so the next editor does not “optimize” it back.
| Pattern | How it bounds work | Cost | Use when |
|---|---|---|---|
| Smaller lists | Pagination / query limit / hard slice | Honest | You can name max items |
| Batch then one child | Loop Over Items → one Execute Sub-workflow per batch | Extra nodes | Mapping + a modest write |
| Wait in the child | Sleep between HTTP | Wall-clock | One vendor, one graph |
| Single-consumer workflow | All producers enqueue; one published graph drains | Ops | Many parents share one low API budget |
| Webhook-triggered children | Each child is a production job under the real cap | Must key writes | You need cross-worker fan-out |
| External queue (Redis / SQS) + one n8n consumer | True serialization | Real infra | Cap is a hard vendor number, not a vibe |
Procedure we use when a parent is already in production and the cap was a lie:
- Freeze new fan-out. Stop adding per-item Execute Sub-workflow calls.
- Count: one production parent from last night → how many child execution rows, how long they overlapped, which worker (if you can see workers — Settings → Workers is Enterprise self-hosted / Cloud-with-n8n-enablement).
- Put a batch size in front of the child call. Start smaller than you want. You are buying a bound, not throughput.
- Move irreversible writes behind an idempotency claim in the child.
- Only if you still need distribution, change the child start path to a published webhook and treat that child as a new production workflow: cap, error workflow, keys, DLQ.
- Re-run the concurrency=1 proof. Child overlap should now match the batch or the worker cap, depending on which pattern you chose.
- Parent no longer uses unbounded per-item Execute Sub-workflow
- Child writes are keyed
- Error workflow on parent and on webhook children
- One named owner for the vendor budget
- Proof screenshot: Executions overlap vs the number you intended
If the child is a helper with no HTTP, stop. You do not have a concurrency problem. You have a tidy-canvas preference.
Worked pacing, as a heuristic you measure — not a vendor SLA. Airtable documents 5 requests per second per base. If each child run makes 2 HTTP calls and you start 50 children at once, you are not “at concurrency 5.” You are at 100 overlapping calls from one process. A batch of 5 children with a 1-second Wait between batches is a starting ceiling. Confirm with headers, not with hope. Other vendors differ; read their page.
| If each child makes | And you allow this many overlapping children | Rough request pile |
|---|---|---|
| 1 HTTP | 5 | 5 in flight |
| 1 HTTP | 50 (per-item, no Wait) | 50 in flight |
| 4 HTTP, no Wait inside the child | 5 | 20 in flight |
| 4 HTTP, no Wait | 50 | 200 in flight |
The last row is how a “rate-limited-safe” parent still earns a lockout. Count HTTP nodes in the child, then multiply by overlapping child executions.
How to measure whether the cap is actually capping children
The Executions tab will lie if you only look at parents. Measure three meters that should agree and will not, until you fix the graph.
| Meter | Where | What “capped” looks like | What “exempt children” looks like |
|---|---|---|---|
| Production / worker concurrency | Executions header, worker status (if licensed) | Running production jobs ≤ limit | Limit holds; child rows still pile up |
| Redis / Bull jobs | Redis list length, worker logs | Jobs in flight ≈ worker count × concurrency | One job while dozens of child executions run |
| Per-worker CPU | Container / host metrics | Load spreads when you add workers | One worker pegged |
| Child overlap | Executions timestamps | Overlap ≤ batch size or ≤ worker cap | Overlap ≈ item count |
Vendor 429 / Retry-After | HTTP output | Rare, bounded | Clusters at parent start time |
Proof run (staging, throwaway workflows, no customer writes):
- Child: Wait 15 seconds, no side effects. Save successful production executions.
- Parent: published webhook. Execute Sub-workflow, Run once for each item, 20 static items. Wait for completion on.
- Instance: concurrency 1. One worker if you are in queue mode.
- Fire the webhook once.
- Record: parent count, child count, max overlapping children, Redis jobs (queue mode), which worker logged the work.
- Repeat the parent with children started via production webhook instead of Execute Sub-workflow.
- Compare rows. If row 5 shows 20 overlapping children and one Redis job, your Execute Sub-workflow path is inline and exempt. If row 6 shows children waiting behind the cap, the webhook path is the one the limiter can see.
Do this on a published production URL. Editor Execute is manual. Manual is exempt. A green canvas click is not a load test.
Hedge: Redis key names and Bull queue labels change across n8n versions. If you cannot find the list, still use execution overlap plus per-worker CPU. Those two meters do not require you to reverse-engineer the broker.
How to tell which worker ran the tree when you do not have Settings → Workers:
- Give each worker process a unique hostname or container name.
- Log
HOSTNAME(or the equivalent) in a Code node on the parent and the child, or read worker stdout timestamps. - Fire one parent. If parent log host and every child log host match, the tree did not move.
- Fire a webhook-triggered child fan-out. Hostnames should diversify if more than one worker is idle and the jobs actually queued.
- File those two log greps next to the Executions screenshot.
If you cannot name containers, you cannot prove distribution. “We have three replicas” is a compose file, not a measurement.
Also capture n8n version: Settings → About, or the image tag. The inline-vs-queued behavior is the claim most likely to change. A proof from 1.x is not automatically a proof on 2.x. Re-run the 20-child pair after a major upgrade, before you add workers “because we scaled.”
- Staging parent/child pair exists and is published
- Proof run used the production webhook URL
- Child overlap number is written on the parent workflow
- Version string recorded
- Hostnames recorded for parent vs children
- Webhook-child control run recorded, if you claim you need distribution
No screenshot, no cap. You have a story.
When webhook children are the right move
Use a production webhook (or another production trigger) on the child when you need the limiter and the worker pool to see the work. That is the pattern n8n staff point at when Execute Sub-workflow pins one worker.
It is not a free upgrade. Each child is now an at-least-once production job. The parent HTTP Request can time out. The child can run twice. You need a key, a 200 on duplicates, and an error workflow someone will read.
| Need | Execute Sub-workflow | Webhook child |
|---|---|---|
| Shared mapping, same worker is fine | Yes | No — extra failure surface |
| Return data into the parent in the same run | Yes (Wait on) | Only with a Wait/resume design |
| Cross-worker fan-out | No, typically | Yes |
| Enroll in production concurrency | No | Yes |
| Simple credentials, one graph | Yes | Maybe not — two published workflows to own |
Procedure for a webhook child that will not double-write:
- Child starts with a Webhook node. Publish it. Use the production URL, not the test URL.
- Claim an idempotency key from the business event before CRM / mail / charge.
- Return 200 on duplicate claims. A 500 invites another delivery.
- Attach the shared error workflow. Test it with an automatic fail, not a canvas click.
- Parent sends only after it has a stable event ID. Do not POST a fresh UUID per attempt.
- Leave worker slots for children. If the parent waits on every HTTP child, cap how many parents fan out or you will queue-deadlock.
- Re-run the measurement table. Children should now appear as production jobs.
If you cannot do steps 2–4 this week, do not switch start paths. Batch the Execute Sub-workflow call instead. A slow correct parent beats a distributed double-charge.
Webhook children also count toward quota. Execute Sub-workflow children do not. Do not switch start paths “to get them onto workers” and then act surprised when plan usage jumps. That is the meter working.
Deadlock procedure if parents wait on webhook children:
- Count worker slots: workers ×
--concurrency(or the env override). - Count how many parents can wait at once (their own trigger volume).
- If waiting parents can fill every slot, children have nowhere to run.
- Fixes, pick one: cap concurrent parents; do not wait (then you need a resume/callback and keys); reserve workers; or go back to Execute Sub-workflow with a batch so children stay inline on the waiting parent.
| Waiting parents | Child jobs needing a slot | Slots | Result |
|---|---|---|---|
| 1 | 5, cap allows 5 | 5 | Fine |
| 5, each waiting on 5 children | 25 | 5 | Queue piles; maybe deadlock if parents never free |
| 5 Execute Sub-workflow parents, inline children | 0 extra jobs | 5 | No deadlock; one or more processes may melt |
Inline is not “safer.” It is a different failure. Choose the failure you can see.
What to skip if you only have a week
Skip a queue-mode migration. Skip adding workers. Skip rewriting every helper into a webhook. You have seven days to stop the unbounded child writes, not to become a platform team.
Do this week:
- Inventory activated parents that call Execute Sub-workflow / Execute Workflow.
- For each, record Mode (all items vs per item), Wait option, and whether the child writes anything irreversible.
- On any per-item child that writes: add a batch ceiling the same day. If you cannot batch, disable the parent until you can.
- Put idempotency on those child writes.
- Attach error workflows on parent and child. Prove one automatic fail.
- Run the concurrency=1 / 20-child proof on staging. Save the screenshot next to the workflow description.
| This week | Next, if still hot | Later |
|---|---|---|
| Name every fan-out | Batch size in production | Webhook children + keys |
| Stop per-item writes | Wait pacing on the vendor path | Single-consumer drain |
| Proof run on staging | Per-worker CPU dashboard | Queue mode, if regular mode is the actual bottleneck |
Do not spend the week debating --concurrency=5 vs 10. That knob never saw the children.
A week that only “tunes workers” is a week the vendor still sees the stampede. Change the graph.
When is this not worth doing yet — and when to hire?
DIY when the child is a helper, item counts are small, and you can finish the inventory plus the proof run in an afternoon. Hire when a production parent already fans writes through Execute Sub-workflow and you cannot tell which worker, which overlap, or which rows already landed twice.
| Situation | DIY | Hire / $500 Automation Audit |
|---|---|---|
| Unpublished prototype, no customer writes | Yes — keep children as helpers | No |
| Production mapping child, no HTTP | Yes — leave it | No |
| Production per-item child that writes CRM / mail / money | Batch and key today | Yes, if you cannot name overlap or recover duplicates |
| Queue mode, one hot worker, idle pool | Measure, then shallow or webhook | Yes, if you were about to buy more workers |
| Cloud plan cap “hit” while children still stampede | Bound the graph | Yes, if the team thinks the plan limit is the governor |
| You only ever tested with editor Execute | Finish a published proof | If you cannot publish a throwaway parent/child pair |
Spurlock Studios will not quote a fake share of instances that “hit this.” Open Executions, count child rows under one parent, and look at overlap. That count is the evidence.
If you DIY, still use the handbook defaults: idempotency before writes, a dead-letter path, an error workflow with an execution link. If you hire, buy blast-radius mapping — which parents fan, which children write, which cap you actually have — not a compose file with more replicas.
Adding workers to an Execute Sub-workflow tree is how you spend money to keep the same pin. Fix the start path or the batch size first.
FAQ
Why doesn’t worker concurrency cap my n8n sub-workflows?
Because production concurrency only counts webhook and trigger starts. Execute Sub-workflow still creates a child execution, but that child is documented as exempt — same class as manual, error, and CLI runs. In queue mode the child typically stays on the parent worker as inline work, so --concurrency never sees a second job.
How do I measure whether worker concurrency is actually capping my n8n sub-workflows?
Set concurrency to 1, publish a parent that fans 20 Execute Sub-workflow children with a Wait, and compare overlapping child executions against Redis job count and per-worker CPU. If 20 children overlap while one job runs, the cap is not on children. Repeat with webhook-triggered children to see a limiter that can actually queue them.
What usually fails first when teams try this?
A per-item Execute Sub-workflow write path that treats --concurrency as a vendor governor. One parent occupies one slot; dozens of children hit the API; other workers idle. Second place: adding workers. Third: switching to webhook children without idempotency and double-writing on retry.
How long does this take to show results?
The staging proof is one published parent/child pair and a single webhook fire — usually hours, not a quarter. Putting a batch ceiling on a production fan-out is the same day. Reconstructing which child writes already landed, if you never stored keys, takes longer. That reconstruction is the cost of the unbounded tree, not the cost of the measurement.
What should I skip if I only have a week?
Skip queue-mode builds and extra workers. Do not skip: inventory of Execute Sub-workflow parents, a batch ceiling on any per-item child that writes, idempotency on those writes, error-workflow attach, and the concurrency=1 proof. Leave helper children that do no HTTP alone.
When is this not worth doing yet?
When nothing is published and nothing writes. The moment an activated parent fans child writes, the exemption is already in play — you just cannot see it on the parent cap. A prototype can over-parallelize in the editor. Production cannot use worker count as a substitute for a bound.
CTA
Parent slots are not child work. Count the children.
Read the Production n8n handbook, then use automation or book the $500 Automation Audit.
What questions does this article answer?
- Why doesn’t worker concurrency cap my n8n sub-workflows?
- Because production concurrency only counts webhook and trigger starts. Execute Sub-workflow still creates a child execution, but that child is documented as exempt — same class as manual, error, and CLI runs. In queue mode the child typically stays on the parent worker as inline work, so `--concurrency` never sees a second job.
- How do I measure whether worker concurrency is actually capping my n8n sub-workflows?
- Set concurrency to 1, publish a parent that fans 20 Execute Sub-workflow children with a Wait, and compare overlapping child executions against Redis job count and per-worker CPU. If 20 children overlap while one job runs, the cap is not on children. Repeat with webhook-triggered children to see a limiter that can actually queue them.
- What usually fails first when teams try this?
- A per-item Execute Sub-workflow write path that treats `--concurrency` as a vendor governor. One parent occupies one slot; dozens of children hit the API; other workers idle. Second place: adding workers. Third: switching to webhook children without idempotency and double-writing on retry.
- How long does this take to show results?
- The staging proof is one published parent/child pair and a single webhook fire — usually hours, not a quarter. Putting a batch ceiling on a production fan-out is the same day. Reconstructing which child writes already landed, if you never stored keys, takes longer. That reconstruction is the cost of the unbounded tree, not the cost of the measurement.
- What should I skip if I only have a week?
- Skip queue-mode builds and extra workers. Do not skip: inventory of Execute Sub-workflow parents, a batch ceiling on any per-item child that writes, idempotency on those writes, error-workflow attach, and the concurrency=1 proof. Leave helper children that do no HTTP alone.
- When is this not worth doing yet?
- When nothing is published and nothing writes. The moment an activated parent fans child writes, the exemption is already in play — you just cannot see it on the parent cap. A prototype can over-parallelize in the editor. Production cannot use worker count as a substitute for a bound.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation When does Continue on Fail hide real API errors in n8n
Continue on Fail hides real API errors when the node fails but the run stays green. Error Workflow never fires; last-valid data often walks into the next write.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.