Skip to main content

Keep your API fast—even at peak

Rate limiting and admission checks that help reduce overload and protect expensive work.

Reliable at peak

Set admission budgets for launch spikes and promos before excess work reaches your handlers.

Lower cloud costs

Fewer moving parts to run and monitor. Spend less on Redis and ops.

Dev‑first adoption

Add via SDK. Keep your stack. Clear limits your customers understand.

UNDERSTAND OVERLOAD · KEEP USEFUL WORK MOVING

Keep useful work moving.

For developers, architects, and managers: 27 narrated lessons on why services overload and how to protect them. No cloud experience required. Start with one request, or jump to any lesson.

Course map Lesson 26 of 27 Browse all lessons

Four chapters, one learning path. Read or listen in any order. ~41 minutes total · Times shown at 1×.

Keep enough capacity to finish work and recover.

  1. 01 A useful result starts with one request
  2. 02 Unfinished work grows when arrivals outrun completions
  3. 03 More waiting space does not make work finish faster
  4. 04 A timeout ends waiting, not necessarily the work
  5. 05 Retained work can exhaust a process’s memory
  6. 06 One failed copy can push its neighbors over capacity
  7. 07 Refuse new work while the application can still recover
  8. 08 Different compute platforms still have finite resources
  9. 09 A shared database failure can spread back into every app
  10. 10 Restore capacity before restoring full traffic
  11. 11 A circuit breaker stops repeating a failing call

Contain one customer's excess without punishing everyone.

  1. 12 One customer’s excess should not ruin another’s day
  2. 13 A rate limit checks one customer’s recent allowance
  3. 14 Copies and languages must share the same intended policy
  4. 15 Fairness helps contain damage without reserving every resource
  5. 16 A legitimate crowd can exceed shared capacity

Reduce work entering a shared dependency when it slows.

  1. 17 Measure the operation whose pressure you want to observe
  2. 18 A shared tracker observes; a guard decides; your app acts
  3. 19 Even the fastest recent observations can become slow
  4. 20 Separate local histories can reopen into the same burst
  5. 21 Replace optional expensive work before refusing essential work
  6. 22 A local breaker learns after calls have already piled up
  7. 23 Fast errors explain why the protections remain complementary
  8. 24 Latency feedback helps without measuring exact spare capacity

Choose, combine, and test the right protections.

  1. 25 Ask three different questions before expensive work
  2. 26 Recovery attempts need limits too You are here
  3. 27 Measure useful service, not how much work you accepted

Read the explanation or use the course map above. Audio is optional.

26 / 27

Recovery attempts need limits too

Bounded retry decision

NoYesstill unsuccessfulLost response orrefusalTime + attempt budgetleft?Stop retryingBackoff + jitterSafe, idempotentattempt
  1. Clients Try again immediately

    Many clients at once

  2. Recovering service LESS CAPACITY

    Extra work delays recovery

  3. Result More attempts, few successes

    Load amplification

Illustrative mechanism. Highlight follows the explanation; not a benchmark.

Read the numbered steps in order. The explanation below follows the narration.

Immediate repeated attempts can amplify pressure.

Bound repeated attempts, spread them out, and preserve operation correctness.

0:00
Read transcript & diagram walkthrough

What you will learn

Bound repeated attempts, spread them out, and preserve operation correctness.

After a refusal or a lost response, a client may try again. A retry is another attempt at the same operation. If many clients repeat immediately, they add load just when the service has less capacity.

Backoff means waiting longer between repeated attempts. Jitter means varying that wait so clients do not all return at the same instant. Bound the number of attempts and the total time available; a request whose deadline has already passed should not create endless background work. Honor the service's guidance where its contract provides it.

A missing response does not prove a write failed. Idempotency means repeated attempts have the same intended effect as one attempt. For example, a payment operation must not charge twice simply because the first response was lost. The application must provide the appropriate operation identity and duplicate handling; a rate limit cannot add those semantics for it.

During recovery, new arrivals and repeated attempts both count toward load. Restore traffic gradually, watch useful completions, and keep admission controls active. Another attempt is useful only if it has a realistic chance of delivering a correct result.

Diagram walkthrough
  1. Immediate repeated attempts can amplify pressure.

    Retries are new load. Immediate repeated attempts can concentrate demand while useful capacity is already reduced.

  2. Backoff waits; jitter spreads; bounds stop endless attempts.

    Backoff spaces attempts; jitter prevents every caller choosing the same return instant. Bound attempts and the total remaining deadline.

  3. A lost response does not prove a write failed.

    A lost response does not prove a write failed. Correct duplicate handling needs operation identity and idempotency semantics supplied by the application.

  4. Count new and repeated work during recovery.

    Include retries at every layer in the load budget. Another attempt is useful only while it can still produce a correct result within its limits.

Content revision: 07c7dd00db17

Download review copy (27 lessons)
Check your understanding & apply it

A payment response was lost. Is immediately repeating the write always safe?

Show the explained answer

No. The first write may have committed. Correct duplicate handling and operation semantics are needed, not just another request.

Integration guidance

Use bounded attempts, total deadlines, backoff and jitter. Define duplicate-write handling and keep refused or expired work from becoming an unbounded background queue.

Where the promise stops

A rate limit is not an idempotency mechanism. A timeout does not prove cancellation or a failed write.

Technical depth & sources

Concepts used here

Implementation detail

Include retries at every layer in the total load budget; nested client/service/database attempts can multiply demand.

Original explanation inspired by Fred Hébert and operational references; no endorsement implied.

AI-generated narration: ElevenLabs Eleven v3, Daniel stock voice. Audio streams only when you start listening.

ElevenLabs

Why teams choose us

Happier customers
Consistent response times when it matters most.
Faster delivery
Ship features, not DIY limiters and brittle ops workarounds.
Clear controls
Per‑tenant and per‑route limits your stakeholders can reason about.

Engineering blog

View all posts →

Why you shouldn’t use Redis as a rate limiter: Part 1 of 2

A tour of the common Redis-based rate limiter implementations — and the correctness and performance traps each one hides.

Auto-Scaling Won’t Save You

The myth of infinite serverless scale — why adding machines doesn’t fix overload, and what to do instead.