APIs, integration & security — in depth

GitHub API Rate Limits for CI Pipelines

Secondary rate limits, not hourly quotas, throttle most CI pipelines.

Staff Writer · · 10 min read
Cover illustration for “GitHub API Rate Limits for CI Pipelines”
Rate Limits and Throttling · October 10, 2026 · 10 min read · 2,281 words

A pipeline runs cleanly for weeks, then one morning it starts failing with rate-limit errors. No code changed. Request volume looks the same as last Tuesday. The team checks the hourly quota, finds plenty of room left, and is stuck, because the hourly ceiling was never the limit that mattered.

The full rate limit landscape: primary ceilings, secondary constraints, and token classes

GitHub runs several rate limit systems at once, and the one that actually throttles a given pipeline depends on the token behind the request, the shape of the traffic, and which endpoint is on the receiving end. The most familiar is the primary hourly ceiling, and it varies by token class. Personal access tokens and OAuth tokens share one standard hourly request ceiling. The default GITHUB_TOKEN that GitHub Actions issues automatically carries a lower hourly limit per repository on standard accounts, and a much higher one on GitHub Enterprise Cloud. GitHub Apps installed on an Enterprise Cloud organization sit above both, with a hourly ceiling well beyond the standard user limit, and that tier is where most high-volume CI pipelines belong.

The hourly number is only the first wall. A separate set of secondary limits governs behavior minute to minute rather than hour to hour, and these are what actually throttle most modern automation. Concurrent requests are capped at a fixed maximum regardless of how much hourly budget remains. A single REST endpoint can take 900 read requests per minute and 180 write requests per minute before it pushes back. Content-creation calls, the ones that create issues, post comments, or modify pull requests, are held to 80 per minute and 500 per hour, a noticeably tighter band than general reads or writes. The Search API is the tightest of all at 10 requests per minute, and it is the most common trap for any tool that scans repositories on a repeating basis.

A point system assigns weight to each request type: most GET, HEAD, and OPTIONS calls cost 1 point, while POST, PATCH, PUT, and DELETE cost 5. GraphQL queries follow the same split, 1 point without mutations and 5 with them. Some endpoints carry point costs GitHub hasn't published. The only reliable way to know a pipeline's real cost is to watch it run and track the headers it gets back.

CI-specific request patterns that push pipelines into secondary limit territory

Secondary limits respond to the shape of traffic, not its size, and three shapes occur constantly in CI: parallel bursts, repeated polling, and heavy use of a single endpoint. Each one can trip a secondary limit while the hourly quota sits mostly untouched.

Concurrency is the clearest case. A matrix build that fans out a dozen or more parallel jobs, each authenticating on its own and calling the same endpoint at roughly the same moment, runs into the concurrent-request ceiling before it ever threatens the hourly budget. The jobs don't need to be making many requests each. They just need to be making them at the same time.

Polling causes a slower version of the same problem. If a job checks build status or pull request state every few seconds, it sends the identical request to the identical endpoint whether anything changed or not, and that steady drip burns through the read-requests-per-minute ceiling, a limit that resets every sixty seconds. A script that feels modest by the hour can still look like a burst by the minute.

Endpoint concentration adds a third failure mode. Tooling built around repository scanning, code analysis, or changelog generation tends to lean hard on the Search API, and that endpoint's 10-requests-per-minute ceiling is an order of magnitude tighter than almost anything else GitHub exposes. A tool that runs fine against the general REST API can stall as soon as it starts routing its calls through search.

AI agents tend to combine all three patterns in one process. An agent traversing repository files, posting comments on pull requests, or polling for the completion of a task can hit the write-per-minute ceiling and the content-creation ceiling independently of the hourly quota, and because agents run in loops, a single bug in the loop logic can burn through a large request budget in seconds.

The Tekton Pipelines-as-Code project ran into exactly this kind of exposure and built a direct fix for it. In PR #3007, dated September 28, 2026, the team added a preflight check that reads the token's remaining core API requests immediately after the permission check completes, before any end-to-end test setup begins. If the token has fewer than 30 core API requests left, the job fails immediately. That's a deliberate trade: an early, cheap failure is accepted in place of a late failure that would have burned through expensive cluster setup only to collapse mid-run anyway.

Three architectural decisions that reduce rate limit exposure before any request is made

Most of the exposure described above can be designed away before a pipeline ever makes its first call. Three decisions do most of the work: which token class the pipeline authenticates with, whether it polls or listens for state changes, and whether it asks GitHub for data it already has.

Token class is the highest-leverage change available. Moving from a personal access token, or from the default GITHUB_TOKEN, to a GitHub App installed on a GitHub Enterprise Cloud organization multiplies the hourly ceiling by three. For any pipeline running at real volume, this single swap buys more headroom than most other optimizations combined.

Webhooks replace polling. Instead of a job asking GitHub "has anything changed yet?" on a fixed interval, GitHub pushes an event payload to a listener URL only when something relevant actually happens. A polling loop that checks state hundreds of times an hour to catch one real change spends hundreds of points to learn what a single webhook delivery would have handed over for free.

Conditional requests close most of the remaining gap. A pipeline that nominally fires off hundreds of requests an hour can end up debiting the rate limit for only a small fraction of that total, as long as it's checking state that mostly stays the same between checks.

Batching stacks on top of both. Pulling several fields back in a single GraphQL query, or running one gh pr view... --json title,body,files call instead of three or four separate REST calls, cuts the total request count and reduces how concentrated those calls get on any one endpoint in a given minute.

Handling rate limit responses at runtime: headers, backoff, and the preflight pattern

Design-time decisions reduce exposure, but a pipeline still needs to behave correctly the moment it gets close to a limit, and that behavior depends on reading the signals GitHub already sends back on every response.

The X-RateLimit-Remaining header should be read on every response, not just when something fails. Once that number drops below a safety threshold, the client should slow its request rate or pause outright, rather than continuing at full speed until GitHub returns a 429. Waiting for the rejection means the pipeline has already lost the request that triggered it and has to recover from a failed state.

When a 429 or 403 does arrive, the Retry-After header tells the client how long to wait before trying again. Ignoring that header in favor of a fixed delay is a common mistake, because a fixed delay has no relationship to when the specific limit that was tripped will actually reset, and a client that guesses wrong can walk straight back into the same wall.

The /rate_limit endpoint gives you a periodic overview of quota across resource families, and it costs nothing against the primary limit. It still counts against secondary limits, so hammering it on a tight interval can cause the exact problem it's meant to help avoid. The response headers on ordinary requests are the more reliable real-time source, with /rate_limit serving as an occasional check. The Tekton pattern, checking remaining requests before committing to expensive setup, is the right shape for any end-to-end suite where a mid-run failure wastes more time than an early one.

Backoff needs randomness built in, not just an increasing delay. If several parallel jobs all get rate-limited at the same moment and all retry using the same fixed backoff schedule, they'll hit the limit again at the exact same millisecond they did the first time, because a shared, deterministic delay keeps them synchronized with each other. Adding random jitter to the backoff interval spreads that retry wave out across time, so the jobs stop arriving in lockstep and the concurrent or secondary limit they tripped the first time doesn't just get tripped again immediately. Scheduling non-urgent jobs for off-peak hours helps the same way, by reducing how much contention piles up on a shared token pool, at no code cost.

The wrong foundation: live API testing in CI

Every mitigation covered so far, batching, caching, backoff, token rotation, reduces how often a pipeline hits a rate limit, but none of them remove the dependency that produces the risk in the first place: the pipeline still needs the live GitHub API to be up, fast, and within quota at the exact moment a test runs.

That dependency carries costs that have nothing to do with rate limits directly. Tests that call the live API are slower than tests that don't, and they're non-deterministic, because network latency, transient 5xx responses, and ordinary flakiness all sit outside the pipeline's control. They can't exercise error conditions without actually triggering real errors against a real account. They can't run offline, and they can't run in an isolated CI environment that has no path to the public internet.

Stateless mocks solve isolation, but they still can't hold state across calls. A stub that always returns the same canned response to GET /repos/:owner/:repo/issues/42 has no memory of anything that happened before it was called, so it can't represent a sequence like creating an issue, fetching it, closing it, and fetching it again to confirm the status changed. That sequence is exactly the kind of workflow most GitHub integrations actually need to prove works.

Mocks also go stale. Every change to the real GitHub API requires someone to go back and manually update the stub to match, and that update rarely happens the moment the API changes. Teams tend to find the drift in production, after the mock has spent weeks confidently returning a response GitHub stopped giving out some time earlier.

The more durable foundation is a simulator that holds state across calls, ships with realistic pre-seeded data, and is checked continuously against the real API so the behavior it returns stays current. Tests built on that foundation reflect how GitHub actually behaves, and they don't spend a single unit of rate limit quota to find out.

A stateful GitHub simulator for rate-limit-free CI

The statefulness is what makes the harder workflows testable. Creating a pull request, checking its status, merging it, and confirming the post-merge state is a sequence, and testing it means the thing standing in for GitHub has to remember what happened at each step along the way, a memory a stateless mock cannot hold.

Continuous verification against the real API closes the drift problem that stateless mocks never solve.

Fault injection turns error-handling tests from an accident into a deliberate exercise. A simulator that can return a 429 on demand, simulate added latency, or inject a 503 lets a team confirm that its backoff logic, the exponential delay with jitter described earlier, actually behaves the way it's supposed to, without needing to provoke a real rate limit against a real account to find out.

The same setup matters even more for AI agents calling the GitHub API, traversing repositories, posting comments, managing pull requests on their own. With a deterministic, resettable environment, you can evaluate that agent behavior over and over, without burning quota and without leaving side effects on a real repository each time something goes wrong.

Putting it together: a rate-limit-safe GitHub integration architecture for CI

None of the individual mitigations above are enough on their own. A pipeline that's safe against rate limits is built in layers, with each layer reducing exposure at a different point in the request lifecycle rather than one fix patched onto an otherwise fragile setup.

The first layer is token strategy. Production CI workloads should authenticate through a GitHub App installed on an Enterprise Cloud organization. Personal access tokens stay reserved for developer scripts, and the default GITHUB_TOKEN stays reserved for lightweight Actions automation, because the standard per-repo limit is already enough there.

The second layer is request architecture. Event-driven state changes should travel through webhooks. Anything checked on a schedule should carry ETags and use conditional requests. Multi-field fetches should go through a single batched GraphQL query. Search API calls should be limited to the workflows that genuinely need them, with explicit rate-gating around them given how tight that endpoint's ceiling is.

The third layer is runtime instrumentation. Rate limit headers get read on every response, not just when something breaks. Before you commit to any expensive end-to-end setup, you run a preflight quota check, following the Tekton pattern. Every retry path uses exponential backoff with jitter.

The fourth layer is test infrastructure. Integration and end-to-end tests run against a stateful local simulator, not the live API. Calls to the live API get reserved for smoke tests that confirm the real service is reachable.

Teams running heavy agent workflows on top of live-API CI, sitting at the overlap of the third and fourth layers, carry the highest rate limit risk of any configuration described here. For those teams, the simulator is the prerequisite a stable pipeline is built on.

More in Rate Limits and Throttling