Skip to main content

Keep your API fast—even at peak

Rate limiting and admission checks that help reduce overload and protect expensive work.

Reliable at peak

Set admission budgets for launch spikes and promos before excess work reaches your handlers.

Lower cloud costs

Fewer moving parts to run and monitor. Spend less on Redis and ops.

Dev‑first adoption

Add via SDK. Keep your stack. Clear limits your customers understand.

UNDERSTAND OVERLOAD · KEEP USEFUL WORK MOVING

Keep useful work moving.

For developers, architects, and managers: 27 narrated lessons on why services overload and how to protect them. No cloud experience required. Start with one request, or jump to any lesson.

Course map Lesson 27 of 27 Browse all lessons

Four chapters, one learning path. Read or listen in any order. ~41 minutes total · Times shown at 1×.

Keep enough capacity to finish work and recover.

  1. 01 A useful result starts with one request
  2. 02 Unfinished work grows when arrivals outrun completions
  3. 03 More waiting space does not make work finish faster
  4. 04 A timeout ends waiting, not necessarily the work
  5. 05 Retained work can exhaust a process’s memory
  6. 06 One failed copy can push its neighbors over capacity
  7. 07 Refuse new work while the application can still recover
  8. 08 Different compute platforms still have finite resources
  9. 09 A shared database failure can spread back into every app
  10. 10 Restore capacity before restoring full traffic
  11. 11 A circuit breaker stops repeating a failing call

Contain one customer's excess without punishing everyone.

  1. 12 One customer’s excess should not ruin another’s day
  2. 13 A rate limit checks one customer’s recent allowance
  3. 14 Copies and languages must share the same intended policy
  4. 15 Fairness helps contain damage without reserving every resource
  5. 16 A legitimate crowd can exceed shared capacity

Reduce work entering a shared dependency when it slows.

  1. 17 Measure the operation whose pressure you want to observe
  2. 18 A shared tracker observes; a guard decides; your app acts
  3. 19 Even the fastest recent observations can become slow
  4. 20 Separate local histories can reopen into the same burst
  5. 21 Replace optional expensive work before refusing essential work
  6. 22 A local breaker learns after calls have already piled up
  7. 23 Fast errors explain why the protections remain complementary
  8. 24 Latency feedback helps without measuring exact spare capacity

Choose, combine, and test the right protections.

  1. 25 Ask three different questions before expensive work
  2. 26 Recovery attempts need limits too
  3. 27 Measure useful service, not how much work you accepted You are here

Read the explanation or use the course map above. Audio is optional.

27 / 27

Measure useful service, not how much work you accepted

Useful service outcomes

Local resource boundsApplications can stillrespondTenant fairnessOne tenant's excess iscontainedShared latencyadmissionLess new work enters aslow dependency
  1. Maya Usable report

    Can prepare for her meeting

  2. Viewer Useful suggestions

    Even a safe cheaper list

  3. Operations team Recoverable service

    Avoid repeated system collapse

Illustrative mechanism. Highlight follows the explanation; not a benchmark.

Read the numbered steps in order. The explanation below follows the narration.

Useful customer outcomes are the goal.

Protect useful customer outcomes, while counting refusal, degradation, and failure honestly.

0:00
Read transcript & diagram walkthrough

What you will learn

Protect useful customer outcomes, while counting refusal, degradation, and failure honestly.

Return to the people using the service. Maya needs a report. A viewer wants something useful to watch. An operations team wants a system it can recover without repeatedly losing every application copy. The business value is useful service delivered, not the largest possible count of accepted requests.

Local resource protection preserves the application's ability to respond. Tenant fairness contains one customer's excess. Shared latency feedback helps avoid adding work to a dependency that is already slow. A cheap, correct fallback can keep a customer moving even when optional features are reduced. Refused or failed requests must still be counted honestly.

These protections do not absorb unlimited traffic. Edge protection means filtering traffic before it reaches the application. A distributed denial-of-service attack is a flood intended to overwhelm availability; specialized network protection helps with that scale. A web application firewall filters configured categories of unwanted web requests. Those defenses complement application-aware decisions.

Measure useful responses, delays, explicit refusals, degraded responses, and failures separately. No layer promises total protection. The goal is to preserve more useful, correct service while containing overload. Next, choose one important customer operation. Set its protections and safe fallback, then test both refusal and recovery.

Diagram walkthrough
  1. Useful customer outcomes are the goal.

    The goal is a useful result for a person, not a large count of accepted requests or internally completed operations.

  2. Three protections contribute different kinds of value.

    Measure correct full responses, useful degraded responses, refusals, and failures separately. A cheap fallback can preserve value without pretending every request succeeded.

  3. Outer defenses address other threats and traffic scales.

    Edge filtering and large-scale network defenses complement application-aware admission. No layer can absorb unlimited traffic.

  4. Count the outcomes customers actually experience.

    Choose one important customer operation. Identify local bounds, tenant policy, and the shared dependency. Decide what can safely degrade; test refusal and recovery.

Content revision: 07c7dd00db17

Download review copy (27 lessons)
Check your understanding & apply it

Accepted requests increased, but useful responses fell. Did availability improve?

Show the explained answer

Not for customers who cannot finish their task. Evaluate useful completed outcomes, latency, and deliberate service reductions separately.

Integration guidance

Define customer-success metrics and distinguish normal success, safe degraded success, policy refusal, transport error, timeout, and resource failure.

Where the promise stops

No layer guarantees total protection. DDoS/WAF defenses and application controls solve complementary problems.

Technical depth & sources

Concepts used here

Implementation detail

Compare useful throughput and recovery behavior under realistic mixed workloads. Treat marketing performance numbers as claims requiring workload-specific evidence.

Original explanation inspired by Fred Hébert and operational references; no endorsement implied.

AI-generated narration: ElevenLabs Eleven v3, Daniel stock voice. Audio streams only when you start listening.

ElevenLabs

Why teams choose us

Happier customers
Consistent response times when it matters most.
Faster delivery
Ship features, not DIY limiters and brittle ops workarounds.
Clear controls
Per‑tenant and per‑route limits your stakeholders can reason about.

Engineering blog

View all posts →

Why you shouldn’t use Redis as a rate limiter: Part 1 of 2

A tour of the common Redis-based rate limiter implementations — and the correctness and performance traps each one hides.

Auto-Scaling Won’t Save You

The myth of infinite serverless scale — why adding machines doesn’t fix overload, and what to do instead.