n8n Queue Mode: Switch When Concurrency Hurts, Not Because It Sounds Pro
Enable n8n queue mode when the main process is the bottleneck — UI freeze, webhook lag, piled executions. Redis and workers add ops you now have to own.
William Spurlock Founder — Spurlock Studios Updated 20 MIN
Enable n8n queue mode when the main process is the bottleneck — UI freezes, webhook latency spikes, executions pile up inside one Node process — not because “queue mode” sounds like a grown-up architecture.
Queue mode separates trigger intake from execution via Redis and workers. It does not fix god workflows, missing idempotency, or a box that is simply too small. Hosting choice (Cloud vs self-hosted) is a different decision — see Self-hosted vs n8n Cloud. This post owns the concurrency threshold: when one process cannot both serve the editor and run production work. Broader spine: Production n8n handbook.
The short answer
- Regular mode runs UI, triggers, and executions in one process. Fine until concurrent production load thrashes the event loop.
- Queue mode (
EXECUTIONS_MODE=queue) has the main instance enqueue work; workers pull from Redis and execute. Scale by adding workers (enable queue mode). - You take on Redis + Postgres + shared encryption keys + worker ops. SQLite is not a production queue-mode database (choose n8n’s database).
- Sub-workflows called via Execute Workflow stay on the same worker as the parent — they are not separate queued jobs. That is the landmine.
- Often the real fix is fewer fat workflows or
N8N_CONCURRENCY_PRODUCTION_LIMITbefore you invent a cluster (control concurrency).
When is the main process the bottleneck?
The main process is the bottleneck when intake, editor, and execution share one event loop and execution wins. Queue mode is the move that takes execution off that loop. It is not the move when Redis, Postgres, a vendor API, or a 40-node graph is the actual constraint.
| Signal | Main is the bottleneck | Something else is |
|---|---|---|
| Editor / REST API | Sluggish only while production executions run | Slow even with zero running jobs |
| Webhook TTFB | Climbs as parallel runs climb | Slow with one execution and a fat Code node |
| Process list | One n8n PID at 90%+ CPU | Postgres or Redis CPU/IO is the hot process |
| Memory | Heap climbs with parallel production runs | One execution’s binary payload blows the box |
| After a concurrency cap | UI recovers; jobs wait FIFO | UI still dies; cap never engaged |
Prove it before you buy Redis:
- Capture peak concurrent production executions for 14–30 days.
- Note editor and webhook latency during that peak, not at 2 a.m.
- Set
N8N_CONCURRENCY_PRODUCTION_LIMITand re-measure. - If the UI recovers and the only remaining pain is “jobs wait,” you have a capacity problem — queue mode or more box.
- If the UI stays dead at a low cap, you have a design or hardware problem. Do not cluster that.
Across 500+ automations, the expensive mistake is treating a drowning main process as a hosting brand problem. It is a process-boundary problem.
What problem queue mode actually solves
| Problem | Regular mode | Queue mode |
|---|---|---|
| Many concurrent production executions | Compete inside one Node process | Distributed across workers |
| UI / API responsiveness under load | Degrades when executions hog the loop | Main can stay lighter; workers burn CPU |
| Horizontal scale-out | Vertical only (bigger box) | Add workers |
| Process isolation on crash | One process dies → everything hurts | Worker crash ≠ editor death (main still up) |
| Independent scale of ingress vs jobs | Same process does both | Webhook processors + workers (optional split) |
Official flow, condensed from n8n’s queue-mode docs:
- Main handles timers and webhooks and creates an execution (does not run it).
- Execution ID goes to Redis (Bull queue).
- A worker picks the job, loads workflow data from the database, runs it.
- Worker writes results; Redis notifies main.
It does not solve: bad retry storms, unbounded fan-out, missing DLQ, or vendor rate limits. Those are workflow design. Horizontal workers will run a duplicate-prone graph faster.
What stays on main after you switch
Queue mode moves production execution off main. It does not empty the process.
| Still on main (typical single-main) | Moved to workers |
|---|---|
| Editor UI and static assets | Production workflow graphs |
| Internal REST API | Writes of execution results |
| Timer / poller / persistent-connection triggers (at-most-once work) | The node work those triggers enqueue |
| Production webhook HTTP, unless you add webhook processors | The execution those webhooks enqueue |
| Execution and binary pruning (leader work in multi-main) | — |
That split is why “we turned on queue mode and main is still pegged” is a common post-cutover complaint. You offloaded the graph, then left every inbound HTTP hit on the same PID. Ingress is a second axis — see webhook processors below.
Manual runs still execute on main unless you set OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=true (queue-mode env vars). Leave that off until workers are healthy; debugging a broken worker through the editor is a miserable first week.
Symptoms that regular mode is drowning
Switch when two or more persist under real traffic (not a one-off import):
- Editor or REST API regularly sluggish while executions run.
- Webhook response times climb; providers start redelivering.
- CPU pegged on the single n8n process at peak; memory climbs with parallel runs.
- You already set
N8N_CONCURRENCY_PRODUCTION_LIMITand still cannot meet peak without starving the UI. - You need independent scale of “receive webhooks” vs “run long jobs.”
If the only symptom is “one workflow is a 40-node monster,” split the workflow first.
Checklist before you call it a main-process bottleneck:
- Peak is production traffic, not a manual replay of last month
- Concurrency limit has been tried and measured
- Hot workflow is not a single Code node doing O(n²) work
- Database is not SQLite on a bursting VPS
- Nobody is running
n8n starttwice on the same box “for HA”
Two of those five unchecked and you are about to cluster a mess.
What Redis and workers add operationally
From n8n’s enable queue mode docs: Redis is the broker; the database persists; workers are separate Node processes that pull jobs. You can run Redis on another machine if every n8n process can reach it.
Ops checklist you now own:
- Redis reachable, authenticated, monitored, backed up per your risk posture
- Postgres (or supported DB) — not SQLite for this topology
- Same
N8N_ENCRYPTION_KEYon main, workers, webhook processors - Worker health checks (
QUEUE_HEALTH_CHECK_ACTIVEif you use/healthz) - Graceful shutdown timeout understood (
N8N_GRACEFUL_SHUTDOWN_TIMEOUT, default 30s) - Version pin matching across main and workers
- Named owner for “queue depth is climbing”
- Unique event-log path per worker if they share a writable disk
Worker concurrency defaults to 10 via --concurrency; n8n recommends 5 or higher. Very low concurrency with many workers can exhaust the DB connection pool (same docs).
export EXECUTIONS_MODE=queue
export QUEUE_BULL_REDIS_HOST=localhost
export N8N_ENCRYPTION_KEY=<same-as-main>
n8n worker --concurrency=5
Redis knobs that bite in production (defaults from the queue-mode variable list):
| Variable | Default | Why it matters |
|---|---|---|
QUEUE_BULL_REDIS_HOST / PORT | localhost / 6379 | Wrong host looks like “queue mode is broken” |
QUEUE_BULL_REDIS_PASSWORD | unset | Open Redis on a public NIC is a credential leak |
QUEUE_BULL_REDIS_TLS | false | Set true when Redis is off-box |
QUEUE_BULL_REDIS_TIMEOUT_THRESHOLD | 10000 ms | How long n8n waits on a dead Redis before exiting |
QUEUE_BULL_REDIS_CLUSTER_NODES | unset | If set, n8n uses a Redis Cluster client and ignores host/port |
QUEUE_WORKER_LOCK_DURATION | 60000 ms | Lease a worker holds on a job |
QUEUE_WORKER_LOCK_RENEW_TIME | 10000 ms | How often the worker renews that lease |
QUEUE_WORKER_STALLED_INTERVAL | 30000 ms | How often to look for stalled jobs (0 = never) |
QUEUE_WORKER_MAX_STALLED_COUNT | 1 | How many times a stalled job is re-processed |
If workers share a filesystem for logs, give each process its own event-log path (per-process event logs). Colliding log files are how you lose the one stack trace that explained last Tuesday.
Bigger box vs fewer god workflows vs queue mode
| Fix | Try when | Stop when |
|---|---|---|
| Raise CPU/RAM | Single process CPU-bound, low concurrency | UI still dies at modest parallel load |
N8N_CONCURRENCY_PRODUCTION_LIMIT | Event-loop thrash; need FIFO backlog | Cap is fine but peak demand needs more machines |
| Split god workflows / move bulk off webhook path | One graph does everything | Graphs are already small; volume is real |
| Queue mode + workers | Sustained parallel production load + ops capacity | Nobody will own Redis/Postgres |
| Webhook processors | Ingress RPS drowns main after workers exist | The graph, not HTTP, is slow |
| Multi-main | You need HA for UI/API/triggers and you have Enterprise | You have not proven workers yet |
Concurrency control in regular mode queues excess production executions FIFO when you set the limit (control concurrency). That is often the right intermediate step. It is also the cheapest proof that the main process — not the vendor API — is what hurts.
Do not skip the middle row because a blog told you “real production uses workers.” Regular mode plus a cap plus smaller graphs is a valid permanent posture. I have shipped that posture on purpose.
Regular-mode concurrency is not queue mode
People mash these together. They are different mechanisms that share one env var.
| Regular mode cap | Queue mode --concurrency | |
|---|---|---|
| What it limits | Production executions on the instance | Jobs one worker may run in parallel |
| Default | Off (unlimited production concurrency) | 10 per worker |
| Env / flag | N8N_CONCURRENCY_PRODUCTION_LIMIT | --concurrency, or the same env if not -1 |
| Applies to | Webhook / trigger production runs | Worker jobs |
| Exempt | Manual, error, CLI, sub-workflows | Same exemptions for the production controller |
n8n is explicit: the production concurrency controller does not apply to manual runs, error executions, CLI starts, or sub-workflow executions (control concurrency). You cannot retry a queued execution; cancel or delete removes it from the queue. On startup, n8n resumes queued work up to the limit and re-enqueues the rest.
In queue mode, N8N_CONCURRENCY_PRODUCTION_LIMIT still matters: if it is set to something other than -1, n8n takes the worker limit from that variable and falls back to --concurrency otherwise. Set one source of truth. Do not set 20 in the env and --concurrency=5 and then argue with the box.
Evaluation test runs use a separate limit from production. Do not tune production workers by watching an evaluation slider.
The sub-workflow concurrency landmine
Official concurrency control applies to production executions started from webhooks/triggers. It does not apply to manual runs, error executions, or sub-workflow executions.
In queue mode, Execute Workflow / sub-workflow calls typically continue on the same worker that picked the parent — they are not automatically re-queued as independent Bull jobs. Community confirmation from n8n staff: sub-workflow nodes behave that way; webhook-triggered children can distribute (forum thread).
Failure mode: you scale to five workers, watch one worker sit at 100% while others idle, because a parent with deep sub-workflow trees pins one process for the whole chain. Queue depth looks healthy. The main process looks fine. You buy a sixth worker. The hot worker stays hot.
Mitigations:
- Prefer webhook-triggered children when you truly need cross-worker fan-out.
- Keep sub-workflow trees shallow on hot paths.
- Measure per-worker CPU, not only queue depth.
- Do not assume
--concurrency=5on three workers means fifteen independent sub-calls. - If a parent fans into twenty Execute Workflow nodes, treat that as a design bug, not a scaling ticket.
| Pattern | Distributes across workers? | Use when |
|---|---|---|
| Parent + Execute Workflow children | No — same worker | Shallow helpers, shared mapping |
| Parent enqueues; children start via webhook | Yes — each child is a job | Real fan-out you can idempotency-key |
| Error workflow | Outside the production cap | Parking failures, not throughput |
Idempotency keys belong on the child, not as a comment on the parent. Queue mode will happily run the unsafe version on five boxes.
Webhook processors (optional second scale axis)
Webhook processors are optional. They scale ingress so the main UI is not eating every HTTP hit (enable queue mode):
- Run
n8n webhookwithEXECUTIONS_MODE=queueand Redis/DB access. - Put a load balancer in front; route
/webhook/*and/webhook-waiting/*(send-and-wait / HITL paths) to the webhook pool. - Keep
/webhook-test/*and editor traffic on main. - Optionally disable production webhooks on main (
N8N_DISABLE_PRODUCTION_MAIN_PROCESS). - Do not put main in the webhook load-balancer pool. n8n warns that this dumps production HTTP back onto the editor process.
Workers still run the graph. Webhook processors solve “too many inbound hits,” not “my Code node is O(n²).”
| Path | Route to |
|---|---|
/webhook/* | Webhook processor pool |
/webhook-waiting/* | Webhook processor pool |
/webhook-test/* | Main only |
| Editor, REST, static | Main only |
If you change N8N_ENDPOINT_WEBHOOK, update the balancer the same day. A leftover /webhook/* rule to main is how you “add processors” and change nothing.
Start a processor (Docker shape from the same docs):
docker run --name n8n-webhook \
-e EXECUTIONS_MODE=queue \
docker.n8n.io/n8nio/n8n webhook
Set WEBHOOK_URL on main to the public URL clients already call. Processors listen on the same default port (5678) as main; isolate them with containers or machines, not hope.
Multi-main is not the first switch
Multi-main is self-hosted Enterprise, not Cloud, and it is high availability for main, not a substitute for workers (enable queue mode). If the main process is the bottleneck because it is executing graphs, workers fix that. If you need a second main so a crash does not take the editor and timers with it, that is a later conversation.
In multi-main, followers run regular work (API, UI, webhooks). The leader also runs at-most-once work: timers, pollers, persistent connections, pruning. Leadership moves if the current leader dies or its event loop locks up.
Requirements n8n lists:
- Every main in queue mode, on Postgres and Redis
- Same n8n version on every main and worker
-
N8N_MULTI_MAIN_SETUP_ENABLED=trueon every main - Load balancer with sticky sessions
- License that actually includes the feature
Leader-key defaults: TTL 10s (N8N_MULTI_MAIN_SETUP_KEY_TTL), check interval 3s (N8N_MULTI_MAIN_SETUP_CHECK_INTERVAL). Do not tune these because a blog said “lower is snappier.” Tune them when you have a failover story and a person watching it.
Viewing running workers in Settings → Workers is also Enterprise (self-hosted, or Cloud Enterprise after you contact n8n). Without that UI, you watch process metrics and Redis depth yourself.
Binary data and large webhook responses
n8n documents that queue mode does not support binary data storage in filesystem mode the way a single process might (handle binary data). For persisted binary under queue mode, the handle-binary-data page tells you to switch N8N_DEFAULT_BINARY_DATA_MODE to database. The queue-mode page points at S3 external storage when you need object storage. External S3/Azure binary modes are self-hosted Enterprise and are not available on n8n Cloud.
Also budget Redis for large Respond to Webhook payloads. The worker returns the response through the queue path. Redis holds the whole body in flight. Default cap is 64 MiB (N8N_WEBHOOK_RESPONSE_RELAY_SIZE_MAX). n8n says to budget about 1.5× that value in Redis memory per response in flight. An oversized reply fails the node unless you offload (large webhook responses).
The same size cap applies to a tool result an MCP Trigger workflow returns from a worker. You cannot offload a tool result; an oversized one surfaces as a tool error that names the limit.
Offload (n8n 2.34.0+):
export N8N_WEBHOOK_RESPONSE_RELAY_SIZE_MAX=64
export N8N_WEBHOOK_RESPONSE_RELAY_OFFLOAD_ENABLED=true
export N8N_DEFAULT_BINARY_DATA_MODE=s3
n8n recommends s3 or azure so main streams chunks. database loads the whole body into main memory and shoves it through Postgres. filesystem needs a shared disk they do not recommend. default (in-memory) has nowhere to offload, so the node still fails above the cap.
Upgrade order is not optional: upgrade every main and webhook instance first, then set the offload flag on workers. An older main will return the storage reference to the client instead of the body.
| Error (from n8n’s table) | Usual cause | Fix |
|---|---|---|
Response too large, names N8N_WEBHOOK_RESPONSE_RELAY_OFFLOAD_ENABLED | Offload off on the worker | Set the flag, or raise the MiB cap |
Response too large, names N8N_DEFAULT_BINARY_DATA_MODE | Mode is in-memory | Switch to a storing mode |
| Too large for the binary-data store | database hit its file cap | Raise N8N_BINARY_DATA_DATABASE_MAX_FILE_SIZE (max 1024 MiB) or leave the DB |
| Stored body could not be read | Main cannot see worker storage | Same bucket/disk/credentials on every process |
Test webhook response size in staging before you cut over. A 70 MiB CSV that worked in regular mode is a Redis-shaped outage in queue mode.
Postgres, SQLite, and the encryption-key failure
Queue mode is a distributed system. Redis brokers jobs; the database persists workflow and execution data. n8n says running that setup on SQLite is not supported (enable queue mode). The database page is softer on regular mode (“not recommended”) and firm on versions: as of July 2026, n8n supports Postgres 17 and 18, plus 16 for compatibility — latest minor inside the major (choose n8n’s database). They do not officially support Aurora, AlloyDB, Cockroach, or Yugabyte. Check that page when you pin; the range moves every November.
n8n Cloud uses SQLite on Starter/Pro (and legacy Enterprise) and Postgres only on Enterprise Scaling plans. That is one more reason Cloud queue mode is an Enterprise conversation, not a toggle.
Encryption-key failure mode: n8n generates N8N_ENCRYPTION_KEY on first start. Workers and webhook processors must share the main key or they cannot decrypt credentials. Symptom is not a clear “wrong key” banner. Symptom is nodes that “cannot use this credential” on workers while the same workflow succeeds on main. Copy the key before you scale out. Rotate it like a secret, not like a Docker comment.
Version-skew failure mode: main on one tag, workers on another. n8n’s queue-mode and external-storage docs both tell you to upgrade components together. Protocol and job-payload mismatches show up as stalled jobs and “it works in the editor.” Pin the image. Deploy main, workers, and processors in one window.
Decision worksheet
Copy into your next ops review:
- Peak concurrent production executions (last 30 days)?
- UI/API latency during that peak — acceptable?
- Already set
N8N_CONCURRENCY_PRODUCTION_LIMIT? Value? - Hours/month available for Redis + worker ops?
- Any workflow with deep Execute Workflow trees on the hot path?
- Webhook RPS vs long-running job count — same process today?
- Binary / Respond-to-Webhook payloads larger than a few megabytes?
- Who gets paged when Redis depth climbs for 15 minutes?
| Answers | Bias |
|---|---|
| (2) fine, (3) unset | Set concurrency limit first |
| (2) bad, (4) near zero | Stay regular or buy managed help; do not DIY cluster |
| (2) bad, (4) solid, (1) high | Queue mode — main is the bottleneck |
| (5) yes | Fix fan-out design before or while switching |
| (6) ingress-heavy, workers already exist | Webhook processors |
| (7) yes | Plan binary mode + relay cap before cutover |
| (8) “we’ll see” | You do not have a cluster. You have a science project. |
If you cannot name the person in (8), you are not ready. Bravery is not an on-call rotation.
Can queue mode hide bad design?
Yes. Horizontal workers will happily run a duplicate-prone graph faster. You get more double-charges per minute.
Before you celebrate queue depth:
- Idempotency on irreversible nodes
- Bound retries / rate-limit pacing
- Error workflow + DLQ path
- Staging proof under parallel load
- No unbounded fan-out from a single webhook
- Execution pruning configured so Postgres does not become the next bottleneck (manage execution data)
Queue mode amplifies whatever you already ship. If the main process is drowning because every Shopify order fans into twelve unrestricted HTTP nodes, workers will drown the vendor instead. That is not an improvement. That is a louder outage.
When to stay on regular mode permanently
Stay regular when:
- Volume is low and the single process is boring under peak.
- No one will monitor Redis/workers.
- You are on n8n Cloud without Enterprise queue mode enabled — Cloud applies plan concurrency limits in regular mode; queue mode on Cloud is Enterprise and requires contacting n8n (Cloud concurrency, pricing).
- The pain is design, not process architecture.
- You have not restored Postgres into a scratch box in the last 90 days.
Regular mode plus a concurrency limit plus smaller workflows is a valid permanent production posture. I will take a boring single process with a named owner over an unwatched Redis every day of the week.
Cloud detail worth repeating: plan limits apply to production executions (webhook/trigger). Manual, error, and sub-workflow runs sit outside that cap — same shape as self-hosted. You do not configure Redis yourself on standard Cloud. If plan concurrency already meets demand, you do not have a queue-mode problem. You have a plan problem or a design problem.
Migration order (regular → queue)
- Postgres + backups proven (restore drill done). Pin a supported major (database versions).
- Redis up with auth, network rules, and TLS if it is off-box.
- Copy
N8N_ENCRYPTION_KEYfrom main. Do not generate a new one on the first worker. - Set
EXECUTIONS_MODE=queueand the shared key on a staging stack first. - Start one worker at
--concurrency=5(or your measured cap). - Run critical workflows under parallel load. Include a fat Respond-to-Webhook path if you have one.
- Confirm sub-workflow hot paths do not pin a single worker unexpectedly.
- Add webhook processors only if ingress still sits on main.
- Production cutover in a maintenance window; watch queue depth and worker CPU for 24–48 hours.
- Keep a rollback: flip
EXECUTIONS_MODEback only if you still have a viable single-process capacity plan.
Skipping staging is how you discover the sub-workflow landmine in front of customers.
If you are migrating databases as part of the cutover, use n8n’s export/import CLI — do not file-copy a SQLite and hope (CLI export).
Health endpoints on each worker, once QUEUE_HEALTH_CHECK_ACTIVE=true:
| Endpoint | Meaning |
|---|---|
/healthz | Process is up |
/healthz/readiness | DB and Redis connections are ready |
Point your platform readiness probe at /healthz/readiness, not at “the container started.” A worker that cannot see Redis will accept traffic and do nothing useful. You can change the health path with N8N_ENDPOINT_HEALTH and the listen port with QUEUE_HEALTH_CHECK_PORT (default 5678 — change it if it collides).
Optional: expose /metrics and scrape workers the same way you scrape main (Prometheus metrics). Queue depth in Redis plus worker CPU plus main event-loop lag is the triangle. One metric lies. Three argue.
What to watch after you switch
| Metric | Healthy shape | Bad shape |
|---|---|---|
| Redis queue depth | Spikes then drains | Climbs without bound |
| Worker CPU | Spread across workers | One worker hot, others idle |
| Main process CPU | Mostly UI/API/webhooks | Still pegged (ingress not offloaded) |
| Execution age (start → finish) | Stable vs baseline | Latency up with no design change |
| DB connections | Under pool max | Exhaustion at low --concurrency + many workers |
| Redis memory | Flat relative to job size | Climbs with Respond-to-Webhook payload size |
| Stalled job count | Near zero | Re-processing loop (lock duration vs job length) |
| Webhook 429 / retries from vendors | Baseline | Climbing — processors missing or main still in the pool |
If one worker is always hot, inspect Execute Workflow depth before you buy more machines.
If main is still pegged after workers are healthy, you did not finish the job. The main process is still the bottleneck — now on ingress, not execution. Add webhook processors or stop putting main in the webhook pool.
If queue depth climbs and every worker is busy, you have real load. Add workers or lower per-workflow work. That is the case queue mode exists for.
If queue depth climbs and workers are idle, you have a broker, auth, or version-skew problem. Do not add machines to a queue nothing is pulling.
FAQ
Does n8n Cloud use queue mode for me?
Not by default. Cloud enforces plan-based concurrency limits for production executions in regular mode. Queue mode on Cloud is available on Enterprise plans if you contact n8n to enable it (Cloud concurrency). You do not configure Redis yourself on standard Cloud.
Do sub-workflows respect worker concurrency?
Sub-workflow executions are exempt from the production concurrency controller. In queue mode they generally run on the same worker as the parent rather than as separate queued jobs, so deep trees can pin one worker while others idle. Webhook-triggered children can distribute; Execute Workflow children usually cannot.
Do I need Postgres for queue mode?
Use a real multi-process database. n8n documents that a distributed queue-mode setup over SQLite is not supported — Redis brokers jobs and the database persists execution data. Plan on Postgres (supported majors per n8n’s database page; as of July 2026 that is 17 and 18, plus 16 for compatibility).
How do webhook processors fit?
Optional processes that accept production webhook HTTP and enqueue executions so main is not the only ingress. They still need queue mode, Redis, and the shared encryption key. Pair them with a load balancer and workers that actually run the workflows. Keep /webhook-test/* on main.
Can queue mode hide bad workflow design?
Yes. It scales execution of whatever you built — including duplicate side effects and retry storms. Fix idempotency, bounds, and DLQ before or as you scale workers. A drowning main process caused by unbounded fan-out becomes a drowning vendor after the switch.
When should I stay on regular mode permanently?
When peak load is fine on one process, when you lack ops capacity for Redis/workers, or when Cloud plan concurrency already meets demand. Prestige is not an availability strategy. Regular mode plus a concurrency cap plus smaller graphs is a legitimate production posture.
CTA
Switch when the main process is the bottleneck — and only if someone will own Redis.
Read the handbook for the rest of the spine, then use automation or book the audit if you want a concurrency threshold review before you stand up a cluster.
What questions does this article answer?
- Does n8n Cloud use queue mode for me?
- Not by default. Cloud enforces plan-based concurrency limits for production executions in regular mode. Queue mode on Cloud is available on Enterprise plans if you contact n8n to enable it ([Cloud concurrency](https://docs.n8n.io/deploy/use-n8n-cloud/understand-concurrency/)). You do not configure Redis yourself on standard Cloud.
- Do sub-workflows respect worker concurrency?
- Sub-workflow executions are exempt from the production concurrency controller. In queue mode they generally run on the same worker as the parent rather than as separate queued jobs, so deep trees can pin one worker while others idle. Webhook-triggered children can distribute; Execute Workflow children usually cannot.
- Do I need Postgres for queue mode?
- Use a real multi-process database. n8n documents that a distributed queue-mode setup over SQLite is not supported — Redis brokers jobs and the database persists execution data. Plan on Postgres (supported majors per [n8n's database page](https://docs.n8n.io/deploy/host-n8n/configure-n8n/choose-n8ns-database/); as of July 2026 that is 17 and 18, plus 16 for compatibility).
- How do webhook processors fit?
- Optional processes that accept production webhook HTTP and enqueue executions so main is not the only ingress. They still need queue mode, Redis, and the shared encryption key. Pair them with a load balancer and workers that actually run the workflows. Keep `/webhook-test/*` on main.
- Can queue mode hide bad workflow design?
- Yes. It scales execution of whatever you built — including duplicate side effects and retry storms. Fix idempotency, bounds, and DLQ before or as you scale workers. A drowning main process caused by unbounded fan-out becomes a drowning vendor after the switch.
- When should I stay on regular mode permanently?
- When peak load is fine on one process, when you lack ops capacity for Redis/workers, or when Cloud plan concurrency already meets demand. Prestige is not an availability strategy. Regular mode plus a concurrency cap plus smaller graphs is a legitimate production posture.
Last reviewed
Automation
Automation After the show is not you at 1 a.m.
Post-show onboarding — thank-you, join path, merch nudge — belongs in a human-gated n8n rail, not your thumb at load-out.
Automation Paperwork that is not the plant
Invoice and PO matching, intake, and support triage in n8n with Metrc fences — the paperwork operators hate, not a menu widget.
Automation Saturday still books — the missed-call rail for trades
A missed-call text-back that routes zip and books a slot beats voicemail and Saturday desk coverage you cannot keep staffed. If a kid is cheaper, say so.
Automation When does Continue on Fail hide real API errors in n8n
Continue on Fail hides real API errors when the node fails but the run stays green. Error Workflow never fires; last-valid data often walks into the next write.
Will's Journal in your inbox.
What I learned this week building for shops, floors, and houses.
You're on the list.
Sign-up failed — try again.
By subscribing, you agree to the Privacy Policy.