
API rate limiting is a rule that caps how many requests each client can make within a set period of time, such as 100 requests per minute. It sounds like a small detail, but it is often the only thing standing between a working app and one that grinds to a halt because a single client got carried away.
This article starts with that failure and works toward the fix. We'll look at why it happens, how a rate limiter decides what to let through, and how the three classic algorithms differ: fixed window, sliding window and token bucket. Then we'll cover what clients should do when they hit the limit, and the mistakes that make rate limiting unfair or useless.
- Rate limiting caps requests per client, per time window, such as 100 per minute per API key.
- It protects your servers, keeps costs predictable, and keeps service fair for everyone.
- Fixed window is simplest, sliding window removes boundary bursts, and token bucket allows friendly short bursts.
- Servers signal the limit with HTTP 429 Too Many Requests, often with a Retry-After header.
- Clients should back off exponentially with jitter, never retry instantly.
- 0:00Intro
- 0:26One busy client can break everyone
- 1:17What rate limiting actually means
- 2:13Protect, save money, stay fair
- 3:01Every request passes the limiter
- 3:54Fixed window counting
- 4:45Bursts at the window boundary
- 5:37The sliding window smooths edges
- 6:33The token bucket
- 7:23Choosing the right algorithm
- 8:05HTTP 429 Too Many Requests
- 8:58Retry with exponential backoff
- 9:51A photo sync hits the limit
- 10:48Where rate limiting goes wrong
- 11:37What you now know
- 12:07Rate limiting is a fairness tool
Why one busy client can slow down everyone
Picture a small online shop whose API comfortably serves a few hundred customers. Then one client starts firing thousands of requests at once. It might be a buggy script or an overeager partner integration. Suddenly pages load slowly, requests time out, and some fail entirely.
The cause is not the total number of users. It is one client taking far more than its fair share. Every request lands on the same API server, so the shopper and the mobile app end up waiting in the same crowded line as the script. Some of their requests get dropped.
To the shopper, the site simply looks broken. They have no idea a script is to blame, and they may never come back. If your hosting scales automatically, it may start extra servers to absorb the flood. That keeps the site alive, but you pay for traffic that gave you nothing.
- Real users see slow pages, timeouts and errors.
- Auto-scaling can quietly raise your bill for unwanted traffic.
What is rate limiting, and why does it matter?
Rate limiting means capping how many requests a client can make in a time window. It has two parts: the cap, which is the maximum number of requests, and the window, which is the stretch of time that number applies to. Tighten either one and the limit feels stricter.
Think of a coffee shop that offers free refills, but only one per customer every ten minutes. It is generous, yet it stops one person from emptying the pot. To apply a rule like that, you need to know who you are counting. Most systems use an API key or a logged-in user, and some fall back to the IP address when nobody is signed in.
The window depends on what you are protecting. A search box might allow a few requests per second, while an expensive report might allow only a handful per day. People often assume limits exist only to stop attackers, but most floods are honest mistakes: a loop that never ends or a retry that fires forever. A limit catches those before they cost real money.
- Protection: load stays within what your servers can handle.
- Cost control: runaway traffic can't produce surprise bills.
- Fairness: heavy users still get served, just not at everyone else's expense.
How a rate limiter checks each request
A rate limiter usually sits at the front door of your system, in an API gateway or middleware. It checks every request before any real work happens. The order is deliberate: the check is cheap, so it runs first, and expensive database work only happens if the request passes.
First, the limiter identifies the client, typically by reading an API key from the request headers. Next, it looks up that client's count, often in a fast shared store like Redis so every server sees the same numbers. Finally, it decides. Under the limit, the request continues and the count goes up. Over it, the limiter responds with an error immediately, which costs almost nothing.
- Identify the client: API key, user ID or IP address.
- Check the counter for the current window.
- Allow and count, or reject right away.
Fixed window counting and the boundary burst
The simplest algorithm splits time into equal blocks, such as each calendar minute, and keeps one counter per client for each block. When a new window opens, the counter starts at zero. Each request adds one, and once the cap is reached, further requests are rejected until the window resets. It works much like a phone plan with a monthly data allowance.
With a limit of 100 per minute, a client that uses all 100 by 12:00:40 is refused until 12:01, when the count starts over. Fixed windows are popular because they store a single number per client, fit in a few lines of logic, and are easy to explain to users.
The catch is the reset. A client that sends 100 requests at 12:00:59 and another 100 at 12:01:00 stays within the rules, yet your server receives 200 requests in about two seconds. That is double the rate you sized for, all at once. For many internal tools a brief spike is harmless, so fixed windows remain a solid choice. Just know the cliff edge exists.
How the sliding window removes the cliff edge
A sliding window stops thinking in calendar minutes. Whenever a request arrives, it asks how many requests this client has made in the last sixty seconds. It is like a speed check that measures your average over the last stretch of road instead of starting over at every mile marker.
Go back to the boundary example. When the second burst arrives at 12:01:00, the 100 requests from 12:00:59 still fall inside the last sixty seconds, so the new burst is refused. Old requests drop out of the count one by one as they age, so capacity never frees up all at once and there is no moment to exploit.
The trade-off is extra bookkeeping. The exact version stores a timestamp for every request. A popular shortcut blends the current and previous window counts into a close estimate. Either way, traffic becomes much smoother.
How the token bucket allows bursts safely
The token bucket, found in many real API gateways, does not count past requests. It hands out permission in advance. Tokens drip into a bucket at a steady rate, up to a maximum size, and each request spends one. If the bucket is empty, the request waits or is rejected. Picture an arcade card that earns one free credit every few seconds but can only hold so many.
This matches how real people use apps. A client that has been quiet builds up a full bucket, so when it suddenly needs to load a page with ten images, it can spend saved tokens all at once. Once those are gone, it is held to the refill rate. You get friendly short bursts and a firm long-term average that nobody can exceed, however hard they try.
Which rate limiting algorithm should you use?
All three algorithms cap requests per client, but they behave differently under pressure. Fixed window is the easiest to build and uses tiny memory, but allows double bursts at resets. Sliding window smooths those edges at the cost of more bookkeeping. Token bucket also stays small in memory and caps bursts at the bucket size, while welcoming short bursts from real users.
There is no single winner. The right choice depends on what your traffic looks like and how much complexity you are willing to manage. A practical rule is to start simple and move up only when you have a reason.
- Fixed window: when short spikes are harmless.
- Sliding window: when boundary bursts cause trouble.
- Token bucket: when users naturally send requests in quick bursts.
What HTTP 429 Too Many Requests means
When a client goes over the limit, the server usually answers with status code 429, Too Many Requests. A good rejection is helpful rather than a slammed door. The response below shows the pattern: a 429 status line, a Retry-After header asking the client to wait thirty seconds, and a plain-language body for whoever is debugging.
A 429 is not the same as a 500 error, which means the server itself broke. It tells the client that each request was fine, but there were too many of them. That distinction shapes the right reaction. Retry-After removes the guesswork: it can be a number of seconds or a specific date, so a well-behaved client knows exactly how long to wait.
HTTP/1.1 429 Too Many Requests
Retry-After: 30
Content-Type: application/json
{"error": "rate limit exceeded"}
Backing off, and avoiding rate limiting mistakes
The worst response to a 429 is retrying instantly, again and again, which only deepens the overload. The right move is exponential backoff. Wait a short moment, say one second, or use Retry-After if the server sent it. If it fails again, double the wait: two seconds, then four, then eight. Add a little randomness, called jitter, so thousands of clients don't retry in sync, set a maximum wait, and give up after a few tries.
Here is how it looks in practice. A newly installed photo backup app has 3,000 photos to upload. The first 50 go out instantly using saved tokens. Then the bucket empties, the server returns 429 with Retry-After, and the app waits thirty seconds before uploading at the refill pace until all 3,000 are done. The server stays healthy and other users never notice. If the app ignored the 429 and retried immediately, one upload would become a self-made traffic storm.
To avoid trouble next time, watch three common mistakes. Clients that hammer back after a 429 turn a small limit into a flood. Counting only by IP address is unfair, because offices, schools and mobile networks share addresses, so one heavy user can block everyone; prefer API keys or user IDs. And hidden limits help nobody: publish your limits and send clear 429 responses so developers can plan around them.
- Back off exponentially with jitter and always respect Retry-After.
- Count by API key or account rather than IP alone.
- Document your limits openly.
Key takeaways
- Rate limiting caps requests per client per time window.
- It protects servers, controls costs and keeps service fair.
- Fixed window is simple but allows bursts at the window boundary.
- Sliding window smooths edges by looking back over a moving period.
- Token bucket allows friendly bursts while keeping a steady average.
- Servers return 429; clients wait and back off before retrying.
Frequently asked questions
What is the difference between rate limiting and throttling?
Both control how fast clients can send requests. In everyday use, rate limiting usually describes the rule itself, while throttling describes slowing requests down rather than rejecting them outright. A token bucket can do either: make a request wait for a token, or reject it.
Should I rate limit by IP address or API key?
Prefer an API key or user account whenever you can. Many people can share one IP address in an office, school or mobile network, so IP-only limits can block innocent users. IP is a reasonable fallback when nobody is signed in.
What should a client do after receiving HTTP 429?
Stop and wait instead of retrying immediately. If the response includes a Retry-After header, wait that long; otherwise use exponential backoff with jitter, a maximum wait, and a limit on the number of retries.
Is the fixed window algorithm good enough?
Often, yes. It is simple, cheap and easy to explain. Its weakness is that a client can send up to double the limit around a window reset, so switch to a sliding window or token bucket if those spikes cause problems.
Why use a token bucket instead of a sliding window?
A token bucket suits traffic that naturally comes in bursts, like loading a page with many images. Clients can spend saved tokens quickly, yet over time they can never exceed the refill rate.