Differentiating Retryable From Non-Retryable API Errors
Distinguish transient failures you can retry from permanent ones that require a fix.

A workflow that calls three APIs in sequence can stall silently when the second call times out and the orchestration code treats that timeout as final. No exception gets thrown, no alert fires, and the downstream steps simply never run, because somewhere in the error handling, a transient failure got filed as a permanent one. This is the actual subject of retry logic: not a list of status codes to memorize, but the discipline of telling a problem that will fix itself apart from a problem that won't, and building the surrounding system so that distinction gets respected every time.
Why the transient/permanent distinction is the only decision that matters
Every retry decision reduces to one question: will sending this request again produce a different outcome? Whether sending the request again will produce a different outcome determines how the failure should be handled, regardless of the status code attached to it, the error message, or whatever pattern a logging dashboard happens to highlight. A transient error, a service that's momentarily overloaded, a gateway that timed out waiting on an upstream dependency, resolves on its own given a little time. The original request was valid; nothing about retrying it changes its shape; the only thing that needs to change is the state of the system receiving it, and that state often repairs itself within seconds.
A permanent error behaves nothing like that. If a request is missing a required field, or references credentials that expired an hour ago, no amount of waiting fixes it.
Getting this distinction wrong carries real cost in both directions. Failing to retry a transient error does the opposite kind of damage: it turns a blip lasting a few hundred milliseconds into a failure the user sees, a transaction that didn't need to fail. Everything else in retry design, the status code tables, the backoff curves, the idempotency keys, exists to serve this one judgment call correctly and consistently.
How HTTP status codes map to retryability
Status codes give that judgment a vocabulary, and most of the time they track the transient/permanent split cleanly enough to act on. Temporal's retry-enhancement recipe adds a code that a lot of standard references skip over: 409 (Conflict), retryable specifically when it signals that a resource is temporarily locked.
On the other side sit the codes that call for a fix, not a resend. A 400 (Bad Request) means a parameter is missing or malformed. A 401 (Unauthorized) means the credential is invalid or expired and needs renewal, not a repeat attempt with the same token. A 403 (Forbidden) means a permission restriction, and retrying does nothing to change it. A 404 (Not Found) says the resource isn't there. A 422 (Unprocessable Entity) means the payload failed validation, so it needs to be corrected before you send it again.
That table is a starting point, and treating it as the whole answer is where a lot of integrations go wrong. Two exceptions demand judgment that no table captures. The first: a 5xx code doesn't automatically mean the failure is transient. A decryption failure, a missing required field that only surfaces deep in server-side processing, a corrupted payload: all of these can come back wearing a 500-range status even though every retry will fail identically. They're permanent errors in disguise, and a retry policy built on status code alone will keep hammering a request that was broken from the moment it was sent.
The second exception runs the other direction. A 404 from a lookup API can be a normal, expected response. Treating that response as a failure, retrying it, alerting on it, throwing an exception in application code, doesn't fix a flaky API. It reveals an integration that misunderstood what the API was telling it.
Why naive retry timing turns a brief outage into a sustained one
Classifying an error correctly only solves half the problem. Getting the timing wrong can turn a transient failure into an outage that outlasts whatever caused it. Retrying immediately after a failure is usually the worst available response, because it throws the exact same load back at a system that just told everyone it couldn't handle the load it already had.
Fixed-interval retries make this worse because they synchronize every failing client onto the same clock. If a server buckles for a moment and every client waits exactly two seconds before trying again, every client arrives back at the same moment, recreating the overload that caused the failure and extending it into a second wave, then a third. Retry timing produces the thundering-herd problem.
Exponential backoff stretches the wait after each failed attempt, so the service gets progressively more breathing room to recover. Jitter, a random offset layered onto the delay, breaks that lockstep by spreading retries across a window of time. A short implementation sketch shows the shape of it:
import random
import time
def retry_with_backoff(fn, max_attempts=4, base_delay=1, max_delay=30):
for attempt in range(max_attempts):
try:
return fn()
except TransientError:
if attempt == max_attempts - 1:
raise
delay = min(max_delay, base_delay * 2 ** attempt)
jittered = random.uniform(0, delay)
time.sleep(jittered)
A few calibration choices determine how well this pattern holds up before it just delays the eventual failure. For jitter, a random value drawn across the entire delay window is the standard default, though AWS guidance describes a more aggressive "full jitter" approach that randomizes the whole delay. If a server sends a Retry-After header, the client should honor it, but a ceiling should still be in place so a server-specified delay can't stall the application indefinitely; Temporal's HTTP retry recipe does exactly this, parsing the header and feeding the resulting delay into a retry policy alongside exponential backoff and timeout settings.
Timing solves the problem of a single client retrying against a single server. It doesn't solve what happens when retry logic exists at several points in the same request path at once, which is a separate and more dangerous failure mode.
How retry logic stacked across application layers compounds into a self-inflicted outage
Retry logic placed at multiple layers of an application doesn't add resilience, it multiplies traffic. The Amazon Builders' Library lays out the arithmetic: retries compounding across several layers of an application stack, each multiplying on the last, can produce a dramatically larger increase in load hitting the downstream database. A brief, recoverable hiccup at the database tier gets amplified by every layer above it that independently decides to retry, until the combined load is what brings the system down.
The fix is architectural rather than tactical: pick exactly one layer of the stack to own retry logic, and disable retries everywhere else in the call path. Five independent retry layers, each reasonable in isolation, compound into something none of them intended.
Circuit breakers add to this: they give the system a way to stop trying once failure crosses a threshold. After enough failures accumulate, the circuit trips and further calls fail immediately. Timeouts need to be set deliberately for connections and requests, because without them one slow downstream service can pin every thread in a pool and exhaust the connections available to the rest of the application, and at that point a retry storm looks indistinguishable from ordinary healthy traffic until the database falls over.
A six-pillar resilience model sequences these concerns in the order they depend on each other: standardized errors make classification possible, classification makes retry decisions possible, safe retries require idempotency, circuit breakers and timeouts bound how much retry traffic the system will tolerate, and observability tells engineers when that budget was set wrong before customers find out the hard way. If you skip a pillar, every pillar that follows inherits its failure mode.
Why idempotency is the precondition for retrying anything that mutates state
Idempotency is the pillar that makes retries safe once classification and timing are both correct, and it only matters for operations that change state. Reading data is inherently safe to retry: asking twice for the same record changes nothing. Creating, updating, or deleting a record is a different matter. If you retry a payment charge without an idempotency guarantee, you charge the customer twice. Retry an order creation and the warehouse ships two packages. A retry that "succeeds" under these conditions is worse than the original failure, because now there's a duplicate to clean up.
An idempotency key solves this by giving the server a way to recognize a replayed request as the same operation. PayPal implements the equivalent idea through the PayPal-Request-Id header on REST POST calls. An IETF Standards Track draft, idempotency-key-header-07, published in October 2025 and expired in April 2026, attempted to generalize an Idempotency-Key header field beyond any single provider's implementation, showing how widely this pattern has been recognized as a gap worth closing at the protocol level.
The key itself has to come from the business operation, not from the request attempt. Generating a fresh random key every time a request gets retried defeats the entire mechanism, because the server has no way to match the retried request back to the original one. Misusing the key creates its own distinct failure: Stripe raises a dedicated IdempotencyError when a key gets reused with parameters that don't match the original request. That's a real conflict in the data, not a retry, and the server is right to refuse it.
Inbound and outbound idempotency solve different problems and shouldn't be treated as interchangeable. Outbound idempotency keys protect against a client accidentally sending the same mutating request twice. Inbound webhook deduplication protects against a provider delivering the same event twice, and Stripe's own guidance is to log every processed event ID and skip any event whose ID has already been logged. One does not substitute for the other.
The HTTP verb itself carries a built-in signal for all of this. GET, HEAD, OPTIONS, PUT, and DELETE are defined as idempotent by the protocol; POST and PATCH are not. That distinction has direct consequences in systems where a duplicate write is expensive: the Kalshi Python SDK retries only GET, HEAD, and OPTIONS requests by default and never automatically retries POST or DELETE, precisely to avoid placing a duplicate order or canceling the wrong position in a financial trading context.
How retry decisions get harder when multiple APIs interact in a single workflow
Idempotency handles a single API retrying its own requests safely. When two APIs in the same workflow each retry on their own, their behaviors combine, and the result can be something neither API's own documentation ever anticipated. The Slack-side consumer has to deduplicate on something stable, like the Stripe payment_intent.id, because handling its own errors correctly isn't enough when the duplication originates upstream.
A Stripe-to-Resend email confirmation flow shows the same problem from a different angle, with an added wrinkle. Scoping a Resend idempotency key to the order reduces duplicate confirmation emails when Stripe retries a webhook delivery, but Resend's idempotency keys don't last forever: they expire after 24 hours. That window is enough to cover most retry storms, but it isn't a permanent guarantee, so an application heading toward a live launch still needs its own outbox pattern or persistent delivery state tracking the email send independently of whatever Resend's key retention happens to cover at the time.
Other providers sit at different points on the spectrum. GitHub issues no automatic webhook retries. PayPal takes the opposite approach, retrying a failed webhook delivery up to 25 times and requiring a 2xx acknowledgment on every attempt before it stops.
Each of these policies sets a floor for what the consuming application has to handle, and that floor is different for every provider. AI agent traffic raises the stakes on this further: agents issue far more calls than a human ever would through a UI, often need their own throttling tier separate from human traffic, and their retry behavior on POST and PATCH endpoints can generate duplicate writes at real scale unless idempotency keys are applied systematically across every call, not patched in after the first incident.
Why Stripe's error taxonomy demands handler logic beyond the HTTP status code
Stripe's own error structure is a clear demonstration of why the HTTP status code was never going to be sufficient on its own. The same 4xx or 5xx status can mean entirely different failures underneath, and each one needs its own handler logic. Stripe's API defines exactly four error types: card_error, invalid_request_error, api_error, and idempotency_error, and each one calls for different handling, not merely a different retry decision. A card_error means the customer's payment method itself was declined and no amount of retrying fixes that. An invalid_request_error points at a malformed call that needs a code fix. An api_error means something went wrong on Stripe's side, and it may be worth a retry under the right backoff policy. An idempotency_error means a key got reused with mismatched parameters, which is a data conflict to resolve, not a network failure to retry past.
The machine-readable code field is what application logic should branch on, not the human-readable message string attached to the error. Messages exist for display to a person and can change wording or get translated without warning; the code is the stable contract between Stripe and the application calling it. An integration that pattern-matches against error message text will break the moment that text changes, even though nothing about the underlying failure condition did. Building retry and error-handling logic around the enumerable, documented code field is what keeps an integration correct as the API evolves around it, and it's the same principle that runs through every section before this one: read what the system is actually telling you, not just the status code that happens to be attached to it.