Spurlock Studios
Contact
Share LinkedIn X
Amber node beads on a dark rail. Thesis: API RATE LIMITS N8N PACE.

Handle API rate limits in n8n by classifying the call, pacing the happy path, then backing off with the vendor’s Retry-After (or documented cool-down) — not by hammering Retry On Fail until the vendor locks you out.

Unlimited retry is how a single burst becomes a multi-hour outage. This post is the production framing we use at Spurlock Studios after 500+ automations. It sits inside the Production n8n handbook.

The short answer

  • 429 means stop and wait — the vendor is throttling you. Read Retry-After when present; do not invent a random backoff that undercuts their cool-down.
  • Retry On Fail is not a rate limiter — it is for transient blips. Bound it. Pair it with Wait / batch pacing on known-slow APIs.
  • Pace before you retry — Loop Over Items (Split In Batches) plus Wait keeps you under the cap so you never enter the 429 spiral.
  • Shed noncritical work — enrichment can fail closed or skip; CRM writes and money moves cannot thrash the same budget.
  • Webhook retries multiply load — a slow 429 loop plus provider redelivery is a stampede. Cap concurrency. Queue mode is a different problem; see when to switch.

What does HTTP 429 actually mean for your n8n design?

RFC 6585 defined 429 Too Many Requests: the client sent too many requests in a given amount of time. The response MAY include Retry-After. Caches MUST NOT store a 429.

RFC 9110 and MDN describe Retry-After as either delay-seconds (Retry-After: 30) or an HTTP-date. Honor the larger of that value and the vendor’s published cool-down.

SignalMeaningDesign response
429 Too Many RequestsYou exceeded a rate or quota windowPause; honor cool-down; reduce concurrency
Retry-After: N (seconds) or HTTP-dateVendor-stated waitWait at least that long before the next attempt
No Retry-AfterCool-down is undocumented or vendor-specificUse the published wait, or jittered exponential backoff
403 with rate-limit headersSome vendors (GitHub) use 403 or 429Treat remaining=0 the same as 429
5xx with retry noiseOften transient infra, not your rate budgetSeparate retry class from 429

Treat 429 as backpressure, not as “try harder.” Trying harder is how you burn the next window too.

n8n surfaces a vendor 429 as The service is receiving too many requests from you in the node output (n8n: Handle rate limits). That string is a hint, not a strategy.

What does n8n document for handling rate limits?

n8n’s own docs give you two integration patterns, plus a third on the HTTP Request node (same page):

PatternWhat it doesUse when
Retry On FailRe-runs the node after a fixed millisecond waitRare blips; wait matches a short vendor window
Loop Over Items + WaitBatches input, pauses between loopsLists, syncs, any hard per-second cap
HTTP Request → BatchingItems per Batch + Batch Interval (ms)One HTTP node, no extra loop graph

Official steps for Retry On Fail:

  1. Open the node → Settings.
  2. Enable Retry On Fail.
  3. Set Max Tries.
  4. Set Wait Between Tries (ms) above the vendor window. Their example: 1 request/second → 1000 ms.

Official steps for the loop:

  1. Put Loop Over Items before the API node. Set Batch Size.
  2. Put Wait after the API node and connect it back to the loop.

Official HTTP Batching:

  1. HTTP Request → Add Option → Batching.
  2. Items per Batch = how many input items ride in one request.
  3. Batch Interval (ms) = pause between those requests.

n8n’s example wait of 1000 ms is correct for a 1/s cap. It is wrong for Airtable’s 30-second 429 cool-down. Do not copy the example number onto a vendor that published a longer lockout.

When is Retry On Fail enough for API rate limits?

Use node-level Retry On Fail when:

  1. Failures are rare and short (one blip, not sustained throttle).
  2. The node is not inside a hot fan-out (hundreds of parallel HTTP calls).
  3. You set a max tries and a delay that matches the vendor, not “infinite.”
  4. The wait you need fits the node’s millisecond field — seconds-long cool-downs belong on a Wait node.
  5. Side effects will not double-write if a late success lands.

When any of those fail, you need pacing or an external queue — not more retries.

SymptomRetry On Fail aloneAdd pacing / Wait
One 503, then greenYesNo
429 every batchNoYes
Shared token across workflowsNoYes, plus inventory
Webhook + slow vendorNoYes, plus fast ack
Cool-down ≥ 10 secondsNoWait node, not ms retry

A GitHub issue on the HTTP Request node records n8n confirming a max of 5 tries and 5000 ms between tries in the editor (n8n#23658). Treat that as a ceiling on the node setting, not as a reason to ignore a 30-second vendor cool-down. If the UI will not hold 30000, the Wait node will.

When do you need Loop Over Items plus Wait?

Pace the happy path for APIs with hard per-second caps or shared budgets across workflows. n8n still ships the node as Loop Over Items; the technical name is Split In Batches (docs).

Minimum pattern:

  1. Collect items (or receive a list).
  2. Loop Over Items — batch size sized to the vendor cap and your parallel branches.
  3. Run the HTTP / vendor node on the batch.
  4. Wait between batches — long enough that peak req/s stays under the limit.
  5. Connect Wait back to the loop. Forgetting the return edge processes one batch and stops.
  6. On 429: Wait using Retry-After (or vendor cool-down), then retry that batch once with a bound.

Wait behavior that matters in production (Wait node):

Wait lengthWhat n8n doesWhy you care
Under 65 secondsProcess stays up; no DB offloadAirtable’s 30s cool-down stays in-process
65 seconds or moreExecution data offloads to the databaseLong cool-downs free the worker
TimezoneServer time, not the workflow timezoneDo not size waits from a local clock assumption

Example Wait expression when the previous HTTP node exposed headers (map the header into the item first):

// Seconds from Retry-After, floor at vendor minimum if header missing
const h = $json.headers?.["retry-after"] ?? $json.headers?.["Retry-After"];
const sec = Number(h);
return Number.isFinite(sec) && sec > 0 ? sec : 30;

Guessing Math.random() * 5 while the vendor wants thirty seconds is how you stay rate-limited.

  • Batch size set from the vendor cap, not from “what looked fast in the editor”
  • Wait sits after the API call and returns to the loop
  • Parallel HTTP lanes on the same token are counted, not ignored
  • 429 path can override the happy-path Wait with Retry-After

Should you honor Retry-After instead of guessing?

Decision list for every production HTTP path:

  1. Capture response status and headers on failure (Error Trigger / Continue On Fail with branch).
  2. If status is 429 and Retry-After is a number → Wait that many seconds.
  3. If Retry-After is an HTTP-date → Wait until that time (or skip and alert if too far out).
  4. If no header → use the vendor-documented cool-down, not folklore.
  5. Retry once (or a small bound). Then stop and shed, or park the item for a human.

Airtable’s Web API documents a concrete case: 5 requests per second per base, plus 50 requests per second across all traffic for a personal access token or service account. Exceeding those returns 429, and Airtable states you must wait 30 seconds before subsequent requests succeed (Airtable rate limits). Their error docs repeat the same 30-second cool-down and the body RATE_LIMIT_REACHED / “Rate limit exceeded. Please try again later” (Airtable errors).

Pace to stay under 5/s. On 429, wait at least 30s. Do not chip away with 2-second retries.

Airtable also ships a different 429: monthly call caps on Free (1,000 calls / workspace / month) and Team (100,000 / month). That error type is PUBLIC_API_BILLING_LIMIT_EXCEEDED. Waiting 30 seconds will not clear a monthly cap (Airtable troubleshooting). Classify the body before you Wait.

Airtable 429What it isWhat works
Per-second rate5/s per base, 50/s per tokenWait ≥30s; shrink batch; shed siblings
Monthly billing capPlan call allotment exhaustedStop. Move the base or wait for reset. Do not retry.

Their support note also says batch create / batch update take up to 10 records per request. Ten records in one call is one request against the 5/s budget. Ten records as ten calls is a self-inflicted 429.

Are vendor rate-limit cool-downs interchangeable?

Copy-pasting an Airtable Wait onto Stripe or GitHub is how you either under-wait (still 429) or over-wait (SLA miss) for no reason.

VendorCap (as of their current docs)429 extrasWhat you honor
Airtable5 req/s per base; 50 req/s per PAT / service account30s cool-down; separate monthly 429Wait ≥30s; batch writes by 10
StripeLive 100 req/s global; sandbox 25 req/s; most endpoints 25 req/sStripe-Rate-Limited-Reason; lock_timeout is also 429Exponential backoff + jitter; read the reason header
GitHub RESTAuthenticated 5,000 req/hour; unauthenticated 60 req/hour403 or 429; x-ratelimit-reset is UTC epochretry-after if present; else wait until reset; else ≥1 minute

Stripe is explicit: treat limits as maximums, watch 429, back off exponentially, and add randomness so retries do not stampede (Stripe rate limits). Live vs sandbox numbers differ on purpose — they discourage load-testing the sandbox as a stand-in for live. A 429 without Stripe-Rate-Limited-Reason may be a lock timeout on one object, not a global rate hit. Serialize mutations on the same PaymentIntent; do not spray parallel updates at it.

GitHub is explicit in the other direction: if x-ratelimit-remaining is 0, do not retry until x-ratelimit-reset (GitHub REST rate limits). Secondary limits may send retry-after. If neither header helps, wait at least one minute, then exponential backoff, then stop. They warn that continuing while limited can get the integration banned.

HeaderWho sends itHow you use it in n8n
Retry-AfterRFC / Airtable-style / GitHub secondaryWait seconds or until HTTP-date
x-ratelimit-remaining + x-ratelimit-resetGitHubIf remaining is 0, Wait until reset epoch
Stripe-Rate-Limited-ReasonStripeBranch: global-rate vs endpoint-concurrency vs lock
None of the aboveMany smaller APIsVendor doc cool-down, then jittered backoff

Do not write one “universal 429 function” that assumes every vendor speaks Airtable.

When should you back off with jitter on a 429?

When the vendor told you how long to wait, wait that long. Jitter is for the case they did not.

AWS’s architecture note is the standard reference: capped exponential backoff without randomness makes every client retry on the same beat — a thundering herd. Add jitter. Full jitter (sleep a random time in [0, min(cap, base * 2^attempt)]) cuts server load the most (Exponential Backoff And Jitter). Stripe says the same thing in their own words: exponential schedule plus randomness.

A Wait expression when Retry-After is absent and the vendor published no cool-down:

// Full jitter, seconds. attempt is 0-based. Cap at 60s.
const attempt = Number($json.attempt ?? 0);
const cap = 60;
const exp = Math.min(cap, 2 ** attempt);
return Math.floor(Math.random() * exp) || 1;

Rules that keep this from becoming folklore:

  1. Prefer Retry-After / vendor cool-down over this formula.
  2. Cap attempts (2–3 after the first failure). Then shed.
  3. Do not jitter below a published minimum (Airtable 30s is a floor, not a suggestion).
  4. Do not apply jitter to a monthly billing 429. That is not a timing problem.
Backoff styleWhen it is honestWhen it is harmful
Honor Retry-AfterHeader presentNever — this is the default
Fixed vendor cool-downDocs state a number (Airtable 30s)Using it on a vendor that resets by epoch
Jittered exponentialNo header, no published waitUndercutting a 30s lockout
Immediate retryNever on 429Always

Should you classify critical vs enrichment before sharing a budget?

Work classExampleOn sustained 429
Critical writeCreate deal, post invoice, route leadBound retry → park → human; pause noncritical siblings
Critical read that gates a writeFetch account before updateSame as write — do not invent the row
EnrichmentFirmographics, AI summary, nice-to-have fieldsFail closed (skip field) or fail open only if product accepts unknown
Bulk backfillNightly syncLower concurrency; extend window; never share burst budget with webhooks

If enrichment and lead routing share one Airtable base and one token, enrichment will starve routing during a scrape. Separate credentials or bases when the business requires it, or shed enrichment first.

Decision list:

  1. Name the token / base / Stripe account this node spends.
  2. Tag the workflow critical or enrichment in the name, not in a comment you will never read at 2am.
  3. Give enrichment its own Wait and a kill switch (disable the workflow).
  4. Never let an enrichment Retry On Fail outrank a lead write on the same 5/s.
  • Every production HTTP node has a work class
  • Enrichment can be disabled without failing the write path
  • Shared-token inventory exists (workflow name, owner, class)
  • Bulk jobs are off the webhook clock

What happens when unlimited retry meets webhook redelivery?

What breaks: a webhook fires; your HTTP node hits 429; Retry On Fail loops; the provider times out and redelivers; now you have N executions all retrying the same throttle.

What it costs: hours of CRM quiet, duplicate side effects if anything eventually succeeds without an idempotency key, and a muted Slack channel full of identical errors.

What you do instead:

  1. Cap retries (small N).
  2. Honor cool-down on a Wait node, not a 1-second Retry On Fail.
  3. Cap workflow concurrency, or put bulk on a single consumer.
  4. Claim an idempotency key before irreversible nodes.
  5. Alert once with an execution deep link, not once per retry.
MultiplierHow it shows upCut it
Node Retry On FailSame execution, same 429, N timesMax tries 2–3
Provider redeliveryNew executions of the same eventFast ack or queue; idempotency
Parallel branchesTwo HTTP lanes on one tokenOne lane per budget
Sibling workflowsEnrichment + routing + backfillShed enrichment; stagger cron

The Wait node staying in-process under 65 seconds means a 30-second Airtable cool-down still occupies the execution. If the webhook provider’s timeout is 10–20 seconds, they will redeliver while you are correctly waiting. Fast-ack the webhook (respond 200, queue the work) or you will do the right Wait and still get stampeded.

How do you rate-limit across multiple n8n workflows?

n8n does not give you a global “Airtable 5/s” governor out of the box. Shared budget patterns:

PatternWhenTradeoff
One “API gateway” workflowMany callers, one vendorExtra hop; clear ownership
External queue (Redis / SQS) + single consumerHigh fan-inReal ops; true serialization
Stagger schedulesCron-heavy estatesEasy; weak under webhook bursts
Separate tokens / basesHard isolation neededCost and admin overhead
N8N_CONCURRENCY_PRODUCTION_LIMITToo many production executions at onceCaps executions, not vendor req/s

n8n’s concurrency control is off by default (-1). Set N8N_CONCURRENCY_PRODUCTION_LIMIT to queue excess production runs FIFO. It applies to webhook/trigger starts, not manual, sub-workflow, error, or CLI runs (n8n concurrency). Useful against a stampede. Useless as a 5/s Airtable governor if each execution still fires five HTTP nodes in parallel.

Checklist for a shared base:

  • Inventory every workflow that hits the vendor
  • Tag critical vs enrichment
  • Cap concurrent executions on the hot paths
  • Put bulk jobs on off-peak schedules
  • One alert owner for 429 storms
  • One consumer if more than two producers share a 5/s cap

Does n8n queue mode replace a vendor rate governor?

Stay inside n8n pacing when volume is moderate and one or two workflows own the vendor.

Add an external queue when:

  • Many producers share one low cap (classic Airtable/base case).
  • You need fair scheduling across clients or brands.
  • Backfills must not starve interactive webhooks.
  • You already run Redis/SQS for other reasons and can put a single consumer in front of the API.

Queue mode (workers + Redis) scales execution concurrency. Worker --concurrency defaults to 10. N8N_CONCURRENCY_PRODUCTION_LIMIT overrides that when it is not -1 (n8n queue mode). That is how many graphs run at once. It is not a per-vendor rate governor.

LeverWhat it limitsWhat it does not limit
Loop Over Items + WaitReq/s inside one executionSibling workflows
HTTP BatchingReq/s inside one HTTP nodeOther nodes, other workflows
N8N_CONCURRENCY_PRODUCTION_LIMITProduction executionsHTTP calls per execution
Queue mode workersJobs per workerVendor budget
External queue + one consumerActual vendor req/sNothing, if you size the consumer

If three workers each run a workflow that hits Airtable at 5/s, you have 15/s and a 30-second lockout. Queue mode made that easier, not safer.

How do you size an n8n batch to the vendor cap?

Airtable at 5 req/s per base (docs):

Design choiceExample settingWhy
Batch size5 items if each item = 1 requestStays at the ceiling only if Wait is ≥1s
Better batch10 records per write request, then WaitUses Airtable’s 10-record batch endpoints
Wait between batches≥1 second (prefer 1.1–1.2s)Leaves headroom for other workflows on the same base
Parallel branches1 HTTP lane on that baseTwo parallel 5/s lanes are 10/s — instant 429
On 429Wait ≥30s, then one bounded retryMatches Airtable’s published cool-down
Shared estatePretend you have 1–2 req/s per workflow until measuredShared fiction beats shared outage

Stripe live at 100 req/s global, 25 req/s on most endpoints (docs):

Design choiceExample settingWhy
Happy-path paceStay under 25/s on a single endpointEndpoint cap is tighter than global
Sandbox testsDo not treat 25/s sandbox as live 100/sStripe says the sandbox is a bad load-test stand-in
Same-object writesSerial, not parallellock_timeout 429 is contention, not rate
On 429Read Stripe-Rate-Limited-Reason; jittered backoffNo Retry-After you can trust

GitHub at 5,000 req/hour authenticated (docs):

Design choiceExample settingWhy
Happy-path pace~1.3 req/s sustained, plus burst roomHourly window, not per-second
HeadersMap x-ratelimit-remaining / resetremaining=0 → Wait until reset
Secondary limitHonor retry-after or wait ≥60sThey will ban a stubborn client
Unauthenticated60/hourDo not scrape public repos without a token

If three workflows share an Airtable base, do not each “use the full 5.” Measure, then assign.

Which HTTP Request settings matter under rate limits?

For the n8n HTTP Request node on throttled vendors (common issues):

  1. Retry On Fail — on, but max tries small (2–3).
  2. Wait Between Tries — only for short windows the field will actually hold. Long cool-downs go to a Wait node that reads Retry-After.
  3. Timeout — long enough for the vendor, short enough that webhook providers do not stack redeliveries forever.
  4. Continue On Fail — only when you have an explicit IF on status next; never to swallow 429 into a fake success.
  5. Batching — Items per Batch + Batch Interval when you want the loop without the extra nodes.
  6. Full Response — on when you need headers. You cannot honor Retry-After if you discarded it.
# Example HTTP Request options (UI equivalents)
Retry On Fail: true
Max Tries: 3
Wait Between Tries: 5000   # ms — editor ceiling; do not pretend this is 30s
Batching: on when the node is the only caller
Items per Batch: sized to vendor
Batch Interval (ms): 1100 for a 5/s cap with headroom

Prefer header-driven Wait over a hardcoded 30s when the API sends Retry-After. Prefer 30s over 5s when the API is Airtable and the header is missing.

Continue On Fail pattern that does not lie:

  1. HTTP Request, Continue On Fail on, Full Response on.
  2. IF: status is 2xx → happy path.
  3. IF: status is 429 → Wait (header or vendor) → one retry → park.
  4. ELSE: error workflow with execution link.

Pagination multiplies the budget. Each page is a request. A 200-record Airtable list at 100 records/page is two requests, not one. Filter server-side. Do not page a whole table to find one row on the webhook path.

How do you monitor 429s and then shed load?

SignalWhereAction threshold
Count of 429 responses / hourError workflow → metrics or log drainPage if above baseline ×3
Mean Wait time insertedCustom metric or execution notesRising Wait = budget pressure
Stripe-Rate-Limited-Reason mixHTTP headersconcurrency vs rate vs lock need different fixes
GitHub x-ratelimit-remainingHTTP headersAlert before 0, not after
Webhook redelivery rateProvider dashboardClimbing with your latency → cut work on the hot path
Monthly Airtable usageWorkspace → Usage429 with billing type → stop retries

One Slack message with deep links beats fifty identical “Rate limit exceeded” lines.

Shed procedure when the storm is already on:

  1. Pause enrichment and bulk sync workflows sharing the token/base.
  2. Leave the critical write path running at reduced concurrency.
  3. Drain in-flight retries; do not Replay All.
  4. Confirm 429 rate drops.
  5. Resume enrichment at half prior batch size.
  6. Schedule the architecture fix (gateway consumer or separate base) within a week.

Shedding is not permanent architecture. It is how you buy the hour to fix architecture.

Replay All after a 429 storm is how you buy a second storm. Replay one execution, confirm cool-down, then drip.

What rate-limit checklist should operators ship this week?

  • Every production HTTP node: max retries set, not unlimited
  • 429 path reads Retry-After or vendor cool-down
  • Full Response on if you need headers
  • Hot lists use Loop Over Items + Wait, or HTTP Batching, sized to the cap
  • Enrichment can be skipped without failing the critical path
  • Shared-token inventory exists
  • 429 storm alert pages a human once, with a mute plan
  • Idempotency on irreversible side effects
  • Webhook path acks fast enough that Wait cannot trigger redelivery
  • Queue mode / concurrency caps are not mistaken for a vendor governor
  • Shed procedure written where on-call can find it
  • Replay All is not the recovery plan

FAQ

Why does unlimited retry make rate limits worse?

Each failed attempt still counts against many vendors’ windows, and aggressive retries keep you pinned at the ceiling. You never drain the cool-down, so every subsequent call fails too. Bound retries and wait the documented interval.

How do I rate-limit across multiple workflows?

n8n will not automatically share a per-base budget across workflows. Inventory callers, pace each hot path, stagger bulk jobs, and for hard caps put a single consumer (gateway workflow or external queue) in front of the API. Queue mode and N8N_CONCURRENCY_PRODUCTION_LIMIT cap executions, not vendor requests per second.

What about Airtable’s 5 requests/second?

Airtable’s Web API is limited to 5 requests per second per base, and 50 requests per second for all traffic using a given personal access token or service account. On exceed you get 429 and must wait 30 seconds before requests succeed again (docs). Design batches under 5/s; on 429 wait ≥30s. A monthly-cap 429 will not clear in 30 seconds.

Should enrichment fail open or closed?

Default fail closed: skip the enrichment field and continue the critical write when the product accepts a thinner record. Fail open (proceed without data) only when a missing enrichment cannot corrupt downstream decisions. Never let enrichment retries starve critical writes on the same budget.

How do bursts from webhook retries interact with limits?

Provider redelivery plus your own Retry On Fail multiplies concurrent calls into the same throttle. Cap retries, ack or queue quickly, use idempotency, and keep bulk/enrichment off the webhook hot path. A correct 30-second Wait still occupies the execution and can trigger redelivery if the provider times out first.

When do I need an external queue?

When many producers share a low vendor cap, when backfills must not starve interactive traffic, or when you need fair multi-tenant scheduling. n8n Wait/batch is enough for a single paced workflow. Shared estates need a real governor. Queue mode is not that governor.

CTA

Pace first, honor the cool-down, then shed — retries are the last tool, not the first.

For the full production spine, keep the handbook open. When you want a rate-limit and backpressure review on your stack, use automation or book the audit.

FAQ

What questions does this article answer?

Why does unlimited retry make rate limits worse?
Each failed attempt still counts against many vendors' windows, and aggressive retries keep you pinned at the ceiling. You never drain the cool-down, so every subsequent call fails too. Bound retries and wait the documented interval.
How do I rate-limit across multiple workflows?
n8n will not automatically share a per-base budget across workflows. Inventory callers, pace each hot path, stagger bulk jobs, and for hard caps put a single consumer (gateway workflow or external queue) in front of the API. Queue mode and `N8N_CONCURRENCY_PRODUCTION_LIMIT` cap executions, not vendor requests per second.
What about Airtable's 5 requests/second?
Airtable's Web API is limited to 5 requests per second per base, and 50 requests per second for all traffic using a given personal access token or service account. On exceed you get 429 and must wait 30 seconds before requests succeed again ([docs](https://airtable.com/developers/web/api/rate-limits)). Design batches under 5/s; on 429 wait ≥30s. A monthly-cap 429 will not clear in 30 seconds.
Should enrichment fail open or closed?
Default fail closed: skip the enrichment field and continue the critical write when the product accepts a thinner record. Fail open (proceed without data) only when a missing enrichment cannot corrupt downstream decisions. Never let enrichment retries starve critical writes on the same budget.
How do bursts from webhook retries interact with limits?
Provider redelivery plus your own Retry On Fail multiplies concurrent calls into the same throttle. Cap retries, ack or queue quickly, use idempotency, and keep bulk/enrichment off the webhook hot path. A correct 30-second Wait still occupies the execution and can trigger redelivery if the provider times out first.
When do I need an external queue?
When many producers share a low vendor cap, when backfills must not starve interactive traffic, or when you need fair multi-tenant scheduling. n8n Wait/batch is enough for a single paced workflow. Shared estates need a real governor. Queue mode is not that governor.
Sources

Last reviewed

More from this lane

Automation

All →
Book the audit