Skip to main content

Keep your API fast—even at peak

Rate limiting and admission checks that help reduce overload and protect expensive work.

Reliable at peak

Set admission budgets for launch spikes and promos before excess work reaches your handlers.

Lower cloud costs

Fewer moving parts to run and monitor. Spend less on Redis and ops.

Dev‑first adoption

Add via SDK. Keep your stack. Clear limits your customers understand.

UNDERSTAND OVERLOAD · KEEP USEFUL WORK MOVING

Keep useful work moving.

For developers, architects, and managers: 27 narrated lessons on why services overload and how to protect them. No cloud experience required. Start with one request, or jump to any lesson.

Course map Lesson 11 of 27 Browse all lessons

Four chapters, one learning path. Read or listen in any order. ~41 minutes total · Times shown at 1×.

Keep enough capacity to finish work and recover.

  1. 01 A useful result starts with one request
  2. 02 Unfinished work grows when arrivals outrun completions
  3. 03 More waiting space does not make work finish faster
  4. 04 A timeout ends waiting, not necessarily the work
  5. 05 Retained work can exhaust a process’s memory
  6. 06 One failed copy can push its neighbors over capacity
  7. 07 Refuse new work while the application can still recover
  8. 08 Different compute platforms still have finite resources
  9. 09 A shared database failure can spread back into every app
  10. 10 Restore capacity before restoring full traffic
  11. 11 A circuit breaker stops repeating a failing call You are here

Contain one customer's excess without punishing everyone.

  1. 12 One customer’s excess should not ruin another’s day
  2. 13 A rate limit checks one customer’s recent allowance
  3. 14 Copies and languages must share the same intended policy
  4. 15 Fairness helps contain damage without reserving every resource
  5. 16 A legitimate crowd can exceed shared capacity

Reduce work entering a shared dependency when it slows.

  1. 17 Measure the operation whose pressure you want to observe
  2. 18 A shared tracker observes; a guard decides; your app acts
  3. 19 Even the fastest recent observations can become slow
  4. 20 Separate local histories can reopen into the same burst
  5. 21 Replace optional expensive work before refusing essential work
  6. 22 A local breaker learns after calls have already piled up
  7. 23 Fast errors explain why the protections remain complementary
  8. 24 Latency feedback helps without measuring exact spare capacity

Choose, combine, and test the right protections.

  1. 25 Ask three different questions before expensive work
  2. 26 Recovery attempts need limits too
  3. 27 Measure useful service, not how much work you accepted

Read the explanation or use the course map above. Audio is optional.

11 / 27

A circuit breaker stops repeating a failing call

Recovery state machine

Trip criterionOpen wait elapsedProbes succeedProbes failCLOSED · allow callsOPEN · refuse locallyHALF OPEN · limitedprobes
  1. Application Requests suggestions

    Caller needs an answer

  2. Suggestions service Can fail or stall

    Separate dependency

  3. Viewer Needs a usable page

    Suggestions are optional

Illustrative mechanism. Highlight follows the explanation; not a benchmark.

Read the numbered steps in order. The explanation below follows the narration.

A recommendations dependency can fail while the application still works.

Closed permits calls; open refuses them; half-open tests recovery with bounded probes.

0:00
Read transcript & diagram walkthrough

What you will learn

Closed permits calls; open refuses them; half-open tests recovery with bounded probes.

Imagine a video website asking another service for suggested videos. That service is a dependency. If its calls repeatedly fail, a circuit breaker can stop the application from making more of those calls.

Closed means calls are permitted. In a simple example, three observed failures within sixty seconds open the breaker. Open means calls are rejected locally, without waiting for the remote answer. The application can show a precomputed list of this week's releases instead.

After a configured wait, half-open allows a small number of test calls, called probes. Successful probes can close the breaker. Failed probes open it again. This is a repeating recovery cycle, not a promise that waiting repairs the service.

Keep three settings separate: the deadline for each call, the window that counts failures, and the time spent open before testing recovery. Modern breakers can also detect slow calls. They do not all wait for complete failure. Local memory and concurrency bounds still protect the caller while outcomes are pending.

Diagram walkthrough
  1. A recommendations dependency can fail while the application still works.

    A circuit breaker changes whether the caller attempts a dependency operation. A fallback can avoid waiting for another failing call.

  2. Closed → failure threshold → open → wait before bounded probes.

    Example trip rule: three observed failures in a 60-second window. Open refuses locally; it does not cancel every call already admitted.

  3. Half-open branches on probe results; failures return to open.

    After the open wait, half-open admits limited probes. Their outcomes determine the next state; passage of time does not repair the dependency.

  4. A deadline, an observation window and an open-state wait are different clocks.

    Keep call deadline, observation window, and open-state wait separate. Modern implementations can include slow-call criteria as well as errors.

Content revision: 07c7dd00db17

Download review copy (27 lessons)
Check your understanding & apply it

Does an expired open-state wait prove the suggestions service recovered?

Show the explained answer

No. It permits a limited recovery test. Failed probes can reopen the breaker; the application can keep showing a cheap list.

Integration guidance

Choose counted failures, observation window, request deadlines, open-state wait and probe count separately. Keep local resource bounds.

Where the promise stops

Three failures in sixty seconds is an illustrative basic policy, not a universal default. Some breakers also use slow calls or resource limits.

Technical depth & sources

Concepts used here

Implementation detail

A sampling window is not a concurrency ceiling. Define failure classification and cancellation; bound calls already admitted before a trip.

Original explanation inspired by Fred Hébert and operational references; no endorsement implied.

AI-generated narration: ElevenLabs Eleven v3, Daniel stock voice. Audio streams only when you start listening.

ElevenLabs

Why teams choose us

Happier customers
Consistent response times when it matters most.
Faster delivery
Ship features, not DIY limiters and brittle ops workarounds.
Clear controls
Per‑tenant and per‑route limits your stakeholders can reason about.

Engineering blog

View all posts →

Why you shouldn’t use Redis as a rate limiter: Part 1 of 2

A tour of the common Redis-based rate limiter implementations — and the correctness and performance traps each one hides.

Auto-Scaling Won’t Save You

The myth of infinite serverless scale — why adding machines doesn’t fix overload, and what to do instead.