Skip to main content

Keep your API fast—even at peak

Rate limiting and admission checks that help reduce overload and protect expensive work.

Reliable at peak

Set admission budgets for launch spikes and promos before excess work reaches your handlers.

Lower cloud costs

Fewer moving parts to run and monitor. Spend less on Redis and ops.

Dev‑first adoption

Add via SDK. Keep your stack. Clear limits your customers understand.

UNDERSTAND OVERLOAD · KEEP USEFUL WORK MOVING

Keep useful work moving.

For developers, architects, and managers: 27 narrated lessons on why services overload and how to protect them. No cloud experience required. Start with one request, or jump to any lesson.

Course map Lesson 15 of 27 Browse all lessons

Four chapters, one learning path. Read or listen in any order. ~41 minutes total · Times shown at 1×.

Keep enough capacity to finish work and recover.

  1. 01 A useful result starts with one request
  2. 02 Unfinished work grows when arrivals outrun completions
  3. 03 More waiting space does not make work finish faster
  4. 04 A timeout ends waiting, not necessarily the work
  5. 05 Retained work can exhaust a process’s memory
  6. 06 One failed copy can push its neighbors over capacity
  7. 07 Refuse new work while the application can still recover
  8. 08 Different compute platforms still have finite resources
  9. 09 A shared database failure can spread back into every app
  10. 10 Restore capacity before restoring full traffic
  11. 11 A circuit breaker stops repeating a failing call

Contain one customer's excess without punishing everyone.

  1. 12 One customer’s excess should not ruin another’s day
  2. 13 A rate limit checks one customer’s recent allowance
  3. 14 Copies and languages must share the same intended policy
  4. 15 Fairness helps contain damage without reserving every resource You are here
  5. 16 A legitimate crowd can exceed shared capacity

Reduce work entering a shared dependency when it slows.

  1. 17 Measure the operation whose pressure you want to observe
  2. 18 A shared tracker observes; a guard decides; your app acts
  3. 19 Even the fastest recent observations can become slow
  4. 20 Separate local histories can reopen into the same burst
  5. 21 Replace optional expensive work before refusing essential work
  6. 22 A local breaker learns after calls have already piled up
  7. 23 Fast errors explain why the protections remain complementary
  8. 24 Latency feedback helps without measuring exact spare capacity

Choose, combine, and test the right protections.

  1. 25 Ask three different questions before expensive work
  2. 26 Recovery attempts need limits too
  3. 27 Measure useful service, not how much work you accepted

Read the explanation or use the course map above. Audio is optional.

15 / 27

Fairness helps contain damage without reserving every resource

Fairness versus resource isolation

Tenant rate boundBounds recent admitteduseTenant concurrencypoolBounds activeexecution slotsSeparate resourcebudgetBounds that resourcescope
  1. Tenant · N Its useful work

    Should not lose every resource

  2. Containment boundary Limit spread

    Strength depends on mechanism

  3. Tenant · O Excessive work

    Contains the noisy neighbor

Illustrative mechanism. Highlight follows the explanation; not a benchmark.

Read the numbered steps in order. The explanation below follows the narration.

A bulkhead contains the spread of trouble.

Tenant fairness is bulkhead-like containment, not a guarantee of reserved resources.

0:00
Read transcript & diagram walkthrough

What you will learn

Tenant fairness is bulkhead-like containment, not a guarantee of reserved resources.

A ship's bulkhead separates compartments so trouble in one does not flood every other compartment. In software, a bulkhead separates resources or work so one part cannot consume everything. Resource isolation means setting aside or bounding resources for a particular part.

A tenant rate limit is bulkhead-like: it contains excessive arrivals from one customer and makes other customers less likely to suffer. That is a valuable fairness boundary. But it does not necessarily reserve a processor, memory, database connections, or active execution slots for Maya.

Imagine two tenants each start ten requests. One tenant's requests finish quickly. The other's requests remain active and retain much more memory. Equal request counts did not create equal resource use. Work weights can improve the allowance, but strict resource separation may require separate pools or additional concurrency bounds.

Choose the strength of the promise carefully. Say that the rate limit contains a tenant's excessive use. Do not say that rate limiting alone makes every tenant fully isolated from every other tenant.

Diagram walkthrough
  1. A bulkhead contains the spread of trouble.

    A software bulkhead separates or bounds resources so trouble in one part does not consume everything in another part.

  2. A rate limit contains arrivals, not every resource.

    A tenant rate limit is bulkhead-like: it contains excessive arrivals. It does not necessarily reserve memory, connections, or execution slots.

  3. Ten quick requests and ten long requests are different.

    Ten short operations and ten long operations have equal counts but different retained-work costs. Weights improve accounting; they do not create strict isolation.

  4. Match the promise to the protection.

    Name the actual guarantee: recent-use fairness, concurrency separation, or a dedicated resource budget. They are different protections.

Content revision: 07c7dd00db17

Download review copy (27 lessons)
Check your understanding & apply it

Does an equal request allowance reserve equal memory for each tenant?

Show the explained answer

No. Duration and per-request cost differ. Strict resource isolation needs additional resource controls.

Integration guidance

Use tenant rate limits for excessive arrivals; add separate bounded pools or concurrency controls when you need stronger isolation.

Where the promise stops

A rate limit alone does not reserve CPU, memory, active slots, or database connections.

Technical depth & sources

Concepts used here

Implementation detail

Distinguish admission accounting from resource scheduling. State which resources are actually bounded or separated.

Original explanation inspired by Fred Hébert and operational references; no endorsement implied.

AI-generated narration: ElevenLabs Eleven v3, Daniel stock voice. Audio streams only when you start listening.

ElevenLabs

Why teams choose us

Happier customers
Consistent response times when it matters most.
Faster delivery
Ship features, not DIY limiters and brittle ops workarounds.
Clear controls
Per‑tenant and per‑route limits your stakeholders can reason about.

Engineering blog

View all posts →

Why you shouldn’t use Redis as a rate limiter: Part 1 of 2

A tour of the common Redis-based rate limiter implementations — and the correctness and performance traps each one hides.

Auto-Scaling Won’t Save You

The myth of infinite serverless scale — why adding machines doesn’t fix overload, and what to do instead.