Skip to main content

Keep your API fast—even at peak

Rate limiting and admission checks that help reduce overload and protect expensive work.

Reliable at peak

Set admission budgets for launch spikes and promos before excess work reaches your handlers.

Lower cloud costs

Fewer moving parts to run and monitor. Spend less on Redis and ops.

Dev‑first adoption

Add via SDK. Keep your stack. Clear limits your customers understand.

UNDERSTAND OVERLOAD · KEEP USEFUL WORK MOVING

Keep useful work moving.

For developers, architects, and managers: 27 narrated lessons on why services overload and how to protect them. No cloud experience required. Start with one request, or jump to any lesson.

Course map Lesson 16 of 27 Browse all lessons

Four chapters, one learning path. Read or listen in any order. ~41 minutes total · Times shown at 1×.

Keep enough capacity to finish work and recover.

  1. 01 A useful result starts with one request
  2. 02 Unfinished work grows when arrivals outrun completions
  3. 03 More waiting space does not make work finish faster
  4. 04 A timeout ends waiting, not necessarily the work
  5. 05 Retained work can exhaust a process’s memory
  6. 06 One failed copy can push its neighbors over capacity
  7. 07 Refuse new work while the application can still recover
  8. 08 Different compute platforms still have finite resources
  9. 09 A shared database failure can spread back into every app
  10. 10 Restore capacity before restoring full traffic
  11. 11 A circuit breaker stops repeating a failing call

Contain one customer's excess without punishing everyone.

  1. 12 One customer’s excess should not ruin another’s day
  2. 13 A rate limit checks one customer’s recent allowance
  3. 14 Copies and languages must share the same intended policy
  4. 15 Fairness helps contain damage without reserving every resource
  5. 16 A legitimate crowd can exceed shared capacity You are here

Reduce work entering a shared dependency when it slows.

  1. 17 Measure the operation whose pressure you want to observe
  2. 18 A shared tracker observes; a guard decides; your app acts
  3. 19 Even the fastest recent observations can become slow
  4. 20 Separate local histories can reopen into the same burst
  5. 21 Replace optional expensive work before refusing essential work
  6. 22 A local breaker learns after calls have already piled up
  7. 23 Fast errors explain why the protections remain complementary
  8. 24 Latency feedback helps without measuring exact spare capacity

Choose, combine, and test the right protections.

  1. 25 Ask three different questions before expensive work
  2. 26 Recovery attempts need limits too
  3. 27 Measure useful service, not how much work you accepted

Read the explanation or use the course map above. Audio is optional.

16 / 27

A legitimate crowd can exceed shared capacity

Legitimate aggregate overload

Viewer A · withinallowanceShared databaseViewer B · withinallowanceMany more legitimateviewersCombined work canexceed capacity
  1. Viewer A · A One suggestion

    Within allowance

  2. Viewer B · B One suggestion

    Within allowance

  3. Thousands more Each within allowance

    Combined demand rises

Illustrative mechanism. Highlight follows the explanation; not a benchmark.

Read the numbered steps in order. The explanation below follows the narration.

A live show brings a legitimate crowd.

All customers can behave reasonably while their combined work overloads the database.

0:00
Read transcript & diagram walkthrough

What you will learn

All customers can behave reasonably while their combined work overloads the database.

Now the service hosts a popular live show's video suggestions. A commercial break brings thousands of viewers at once. Aggregate demand means their combined demand, not what any one viewer sends.

Each viewer requests one suggestion and stays within the individual allowance. Every fairness check can pass. But all those accepted requests still reach the same database. If combined work exceeds the database's capacity, unfinished work accumulates even though nobody broke a rule.

More application copies and local memory protection can keep the application layer alive. They do not automatically increase the database's capacity. A fixed database-wide rate cap can be useful, but its safe value depends on request cost and available capacity. One simple lookup is not the same work as a large report.

This is a different problem from Orchid's bulk job. We are no longer asking which tenant exceeded an allowance. We are asking whether the shared resource can handle more work now. Passing the fairness check alone does not answer that question.

Diagram walkthrough
  1. A live show brings a legitimate crowd.

    A live-event spike can contain only reasonable individual requests. Aggregate demand is their combined work.

  2. Every fairness check can pass while aggregate work is too high.

    Every per-tenant check can pass while the shared database receives more work than it can finish.

  3. Adding application copies does not add database capacity.

    Application scaling and local guards do not add database capacity. A fixed aggregate cap also depends on request cost and available capacity.

  4. Fairness and current shared capacity are different questions.

    Fairness asks who exceeded a policy. Shared-resource admission asks whether adding work looks safe now. One answer cannot substitute for the other.

Content revision: 07c7dd00db17

Download review copy (27 lessons)
Check your understanding & apply it

Every viewer is within their allowance. Can the shared database still overload?

Show the explained answer

Yes. The sum of individually reasonable use can exceed the database’s available capacity.

Integration guidance

Keep tenant fairness, but also bound and observe work sent to shared dependencies. Test changes in workload cost and available capacity.

Where the promise stops

A fixed shared-resource cap can be a useful guardrail; it is not automatically well-sized for every workload mix.

Technical depth & sources

Concepts used here

Implementation detail

A user-count spike, costlier queries, or reduced database capacity can all make the same per-tenant policy insufficient.

Original explanation inspired by Fred Hébert and operational references; no endorsement implied.

AI-generated narration: ElevenLabs Eleven v3, Daniel stock voice. Audio streams only when you start listening.

ElevenLabs

Why teams choose us

Happier customers
Consistent response times when it matters most.
Faster delivery
Ship features, not DIY limiters and brittle ops workarounds.
Clear controls
Per‑tenant and per‑route limits your stakeholders can reason about.

Engineering blog

View all posts →

Why you shouldn’t use Redis as a rate limiter: Part 1 of 2

A tour of the common Redis-based rate limiter implementations — and the correctness and performance traps each one hides.

Auto-Scaling Won’t Save You

The myth of infinite serverless scale — why adding machines doesn’t fix overload, and what to do instead.