Stripe Rate Limit Behavior Across API Versions
Stripe enforces five separate rate limits that operate simultaneously in production.

Stripe's rate limiting is a layered system of distinct limits, measured at different scopes and enforced at different points in the request lifecycle, all operating at once against the same account. An engineer who builds retry logic around a single requests-per-second figure has modeled maybe a quarter of the actual system, and the rest of it appears, unannounced, in production. Collapsing the system down to one number is the single most common reason integrations hit unexpected 429 responses after they've already shipped. The Stripe-Rate-Limited-Reason header, which Stripe attaches to rejected requests, carries one of five distinct values: global-rate, endpoint-rate, global-concurrency, endpoint-concurrency, or resource-specific. Five categories of rejection, each tied to a separate enforcement point, confirm that these are not five names for the same limit. They are five different limits, and an integration has to account for all of them.
The four enforcement mechanisms Stripe runs simultaneously
Stripe's engineering team has documented four distinct limiters that run in production at the same time, and only two of them count a client's own requests per second. The request rate limiter is the one most engineers already know about: it restricts each account to a set number of requests per second and functions as the primary control on overall traffic volume. Stripe's engineering blog notes that this limiter rejects a large share of requests in test mode specifically, usually because a script got out of hand without anyone intending it to. The concurrent requests limiter works on a different axis. Rather than counting requests over a one-second window, it caps how many requests can be in flight at any given moment, and it targets CPU-intensive calls such as list requests and requests with expansions, where the cost to Stripe's infrastructure isn't proportional to request count alone.
The other two mechanisms don't look at the client's behavior. The fleet usage load shedder reserves a portion of Stripe's computing capacity for critical API operations, and when that reserved capacity comes under threat, it rejects non-critical requests, a decision driven by the state of Stripe's fleet. The worker utilization load shedder serves as the last line of defense: it sheds non-critical traffic when servers are overloaded and restores that traffic after a delay once conditions recover. Taken together, these four mechanisms mean a client that has correctly throttled itself to stay under its per-second budget can still receive a 429 because a load shedder fired during an incident that had nothing to do with that client's own traffic. Handling 429s correctly requires reading the reason behind the rejection, not just the status code.
The specific numeric limits by resource and environment
The global limits set the floor. Live mode gives you 100 requests per second per account, but sandbox gives you only 25, a quarter of that budget. Individual endpoints default to 25 requests per second unless a tighter limit has been set for that resource, and several resources sit well below the default.
The Payment Intents API caps update requests per PaymentIntent object on an hourly basis, a per-object ceiling that sits entirely apart from the per-second global limit, and you can hit it without the account ever approaching its overall budget. The Subscriptions API scopes its limits to a single subscription object, not to the account as a whole: you get 10 new invoices per subscription per minute, 20 per subscription per day, and 200 quantity updates per subscription per hour. The Files API allows you 20 read requests per second and 20 write requests per second. The Payouts API gives you 15 create requests per second, with a concurrency cap of 30 requests in flight at once. Connect account creation combines the v1 and v2 Accounts endpoints, and you get 30 accounts per second in live mode but only 5 per second in sandbox. The Search API gives you 20 read requests per second, and Stripe's documentation notes that if you need data-intensive analytics, Sigma or Data Pipeline suit it better than repeated Search calls. Issuing card creation limits don't follow a fixed number; they vary by the issuing account's country and industry.
Meter Events, used for billing usage reporting, carry their own budget of 1,000 calls per second in live mode, a figure entirely separate from the 100-per-second global cap. In sandbox, and for connected accounts, Meter Events calls stop drawing from that separate budget, so you count them against the global limit instead. Concurrency limits add a further wrinkle: they don't reset on the one-second cadence that governs per-second rates. They track how many requests are simultaneously active at any instant, which is a different measurement than counting requests over a rolling window.
The sandbox-live testing dead zone
A developer who throttles a bulk job to clear sandbox's 25-requests-per-second ceiling has validated the code against a quarter of the pressure that live mode will apply. Passing sandbox under those conditions says nothing about whether the same job will clear live mode's 100-per-second limit, because the two environments impose genuinely different loads. Sandbox overuse is listed among the common causes of 429 errors precisely because its budget is so much tighter than production's, so a script that trips sandbox limits may be nowhere near the live-mode ceiling, and a job that squeaks past sandbox gives no real assurance it will behave the same way once it moves to live traffic.
Connect account creation makes the gap concrete. Sandbox allows 5 accounts per second; live mode allows 30, a sixfold difference for that single endpoint. A platform onboarding Connect accounts in bulk could pass every sandbox test at a rate that would be unremarkable in production, or could trigger sandbox throttling on a workload that live mode would absorb without strain. Either direction of error is possible, and neither can be ruled out by sandbox testing alone.
The gap compounds once a deployment scales horizontally. Rate limits apply per account, not per API key, so every worker process that calls the API under the same account draws from the same shared budget. A fleet of concurrent workers, each individually well-behaved and each staying comfortably under what it believes is its own allowance, can together exceed the account's single shared ceiling and produce a 429 storm with no single worker at fault. Stripe's support documentation addresses this directly by advising that background jobs be limited to a subset of the account's maximum rate, so that time-sensitive live charges have headroom to succeed alongside bulk traffic.
What the response headers reveal and omit
Stripe's rate-limit headers are built to explain a rejection after the fact. A 429 response carries a Stripe-Rate-Limited-Reason header with one of five values: global-rate, endpoint-rate, global-concurrency, endpoint-concurrency, or resource-specific. A 429 that arrives without that header signals a different kind of rejection. It signals a lock timeout instead, which has its own distinct cause and its own distinct fix, covered in the next section.
Stripe also sets a Stripe-Should-Retry header on some responses: true when Stripe has grounds to believe the request is safe to retry, false when it isn't, and the header is left off entirely when Stripe can't determine retryability one way or the other. What's absent from the header set carries just as much weight as what's present. Stripe does not emit X-RateLimit-Limit, X-RateLimit-Remaining, or Retry-After, the headers that many other major APIs use to let a client see its remaining budget before it runs out. There's no counter a client can poll to back off preemptively. The first signal a client gets that it's approaching a limit is the rejection itself.
That puts the burden of pacing squarely on the developer. Throttling has to be built proactively on the client side, or backoff has to be built reactively to fire the moment a rate-limit error actually arrives, because the API gives no advance warning either way. Developers need their own backoff wrapper to cover that case. The specific value in Stripe-Rate-Limited-Reason should drive how that backoff is scoped: a global-rate rejection calls for slowing every request the account makes, while an endpoint-rate rejection only requires throttling calls to that one endpoint, leaving the rest of the account's traffic free to proceed at normal speed.
Lock timeouts: the 429 that is not a rate limit
A lock-timeout rejection and a rate-limit rejection both arrive under the same HTTP status code, 429, but they come from different mechanisms, so you need to handle them differently. If you treat them as the same error, your retry logic will make the underlying problem worse. Stripe's documentation specifies that a lock-timeout error carries the code lock_timeout and a message telling you that another API request or Stripe process is currently accessing the same object. Critically, it does not carry a Stripe-Rate-Limited-Reason header, which is the clearest signal available for telling the two error types apart.
The cause is concurrency control, not traffic volume. Stripe locks certain objects during some operations so that two concurrent workloads can't produce inconsistent results on the same record, and a lock timeout fires when an incoming request can't acquire a lock that's already held by another process. The fix depends on frequency: an occasional lock timeout on a given object can usually just be retried, but frequent lock timeouts hitting the same object point to a design problem, and the real fix is to serialize requests to that object or reduce the concurrency hitting it directly, rather than to keep retrying against contention that won't resolve on its own.
The distinction matters for idempotency as well. A request that gets rejected with a rate-limit 429 and is then retried with the same idempotency key can produce a different result than the original attempt would have, because Stripe's rate limiters run before the idempotency layer ever sees the request. Lock timeouts don't carry that same risk, because the object-locking behavior occurs at a different point in request handling. A simulator or test harness that doesn't reproduce this distinction, treating every 429 as interchangeable, will validate retry logic that behaves correctly in testing and incorrectly in production.
How API versioning changes the rate-limit surface
Stripe runs two API namespaces, v1 and v2, that differ in consistency model, idempotency behavior, and request format, and each of those differences changes what correct 429 handling looks like depending on which namespace a client is calling. The v1 API uses form-encoded request bodies, and its top-level list endpoints are immediately consistent at the cost of somewhat higher latency. The v2 API uses JSON for both requests and responses, and by default its lists are eventually consistent, so you trade a brief consistency window for lower latency.
That split carries directly into how retries behave after a rate-limit rejection. In v1, retrying a request with an Idempotency-Key returns the previously stored result from the original attempt. In v2, retrying a failed request with an Idempotency-Key reissues the request without producing additional side effects. A retry fired after a 429 therefore has different semantics depending on which namespace issued the original call, and retry logic built against one namespace's behavior will not transfer cleanly to the other.
Versioning adds a second layer of complexity on top of the namespace split. Starting with the 2024-09-30.acacia release, Stripe moved to a cadence of monthly releases that contain no breaking changes, paired with major releases twice a year that can include breaking changes. Stripe pins every account to a specific API version by date, so when a breaking change ships, it applies a transformer chain that downgrades current responses back down to the format that pinned version expects. If a test environment or simulator doesn't replicate that transformer chain, its responses will diverge from what an account pinned to an older version actually receives, so rate-limit testing done against an unversioned or mis-versioned simulation isn't testing the real behavior. Webhook endpoints compound the issue further: each one either carries its own specific API version or falls back to the account's default version, and for statically typed SDKs in.NET, Java, and Go, the webhook version has to match the version the SDK was generated against. A mismatch causes deserialization to fail outright, throwing an exception in.NET and Go, or returning an empty Optional in Java.
The common failure patterns that produce production 429 storms
Most of the 429 storms that hit production systems trace back to a small number of structural patterns, and each one exploits a different layer of the enforcement system described above. Bulk operations run in a tight loop, such as importing or migrating records without any client-side throttling, saturate the global per-second budget directly. Stripe's support documentation advises limiting background jobs to a subset of the account's maximum rate so that live, time-sensitive charges retain headroom. PaymentIntent retry loops create a subtler version of the same failure: repeatedly updating a single PaymentIntent object can hit its 1,000-updates-per-hour cap long before the account comes anywhere near the 100-per-second global ceiling, and that per-object limit stays invisible to engineers who are only watching the global request counter. Webhook fanout produces a third pattern, where a webhook handler processes its payload synchronously by firing off a batch of downstream API calls, multiplying request volume at the exact moment the system is already under load from the event that triggered the webhook, and each subsequent retry adds to the pile when exponential backoff isn't built into the retry path. Horizontal scaling without a coordinating throttle produces a fourth pattern, rooted in the fact that rate limits apply per account rather than per API key: adding more worker processes multiplies the account's total request rate even though no individual worker is doing anything wrong or is even aware the other workers exist. Subscription billing spikes round out the pattern set: because the per-subscription invoice limits of 10 per minute and 20 per day apply to each subscription object individually, a billing platform generating invoices across a large number of subscriptions at once can trip per-object caps on specific subscriptions even while the account's overall request rate stays well within the global limit the whole time.
