Skip to main content

Keep your API fast—even at peak

Rate limiting and admission checks that help reduce overload and protect expensive work.

Reliable at peak

Set admission budgets for launch spikes and promos before excess work reaches your handlers.

Lower cloud costs

Fewer moving parts to run and monitor. Spend less on Redis and ops.

Dev‑first adoption

Add via SDK. Keep your stack. Clear limits your customers understand.

UNDERSTAND OVERLOAD · KEEP USEFUL WORK MOVING

Keep useful work moving.

For developers, architects, and managers: 27 narrated lessons on why services overload and how to protect them. No cloud experience required. Start with one request, or jump to any lesson.

Course map Lesson 24 of 27 Browse all lessons

Four chapters, one learning path. Read or listen in any order. ~41 minutes total · Times shown at 1×.

Keep enough capacity to finish work and recover.

  1. 01 A useful result starts with one request
  2. 02 Unfinished work grows when arrivals outrun completions
  3. 03 More waiting space does not make work finish faster
  4. 04 A timeout ends waiting, not necessarily the work
  5. 05 Retained work can exhaust a process’s memory
  6. 06 One failed copy can push its neighbors over capacity
  7. 07 Refuse new work while the application can still recover
  8. 08 Different compute platforms still have finite resources
  9. 09 A shared database failure can spread back into every app
  10. 10 Restore capacity before restoring full traffic
  11. 11 A circuit breaker stops repeating a failing call

Contain one customer's excess without punishing everyone.

  1. 12 One customer’s excess should not ruin another’s day
  2. 13 A rate limit checks one customer’s recent allowance
  3. 14 Copies and languages must share the same intended policy
  4. 15 Fairness helps contain damage without reserving every resource
  5. 16 A legitimate crowd can exceed shared capacity

Reduce work entering a shared dependency when it slows.

  1. 17 Measure the operation whose pressure you want to observe
  2. 18 A shared tracker observes; a guard decides; your app acts
  3. 19 Even the fastest recent observations can become slow
  4. 20 Separate local histories can reopen into the same burst
  5. 21 Replace optional expensive work before refusing essential work
  6. 22 A local breaker learns after calls have already piled up
  7. 23 Fast errors explain why the protections remain complementary
  8. 24 Latency feedback helps without measuring exact spare capacity You are here

Choose, combine, and test the right protections.

  1. 25 Ask three different questions before expensive work
  2. 26 Recovery attempts need limits too
  3. 27 Measure useful service, not how much work you accepted

Read the explanation or use the course map above. Audio is optional.

24 / 27

Latency feedback helps without measuring exact spare capacity

Little’s law worked example

100 completions / s ×0.1 s average10 operations inflight on average100 completions / s ×0.5 s average50 operations inflight on average
  1. Throughput 100 operations / s

    Same measurement boundary

  2. Average time 0.1 s

    Not a minimum

  3. Average in flight 10 operations

    100 × 0.1

Illustrative mechanism. Highlight follows the explanation; not a benchmark.

Read the numbered steps in order. The explanation below follows the narration.

Stable example: throughput × average time = average in flight.

Latency is useful feedback, not an exact queue or spare-capacity measurement.

0:00
Read transcript & diagram walkthrough

What you will learn

Latency is useful feedback, not an exact queue or spare-capacity measurement.

The latency signal is useful because longer operations can keep more work unfinished. Consider a system completing one hundred operations per second. In a stable period, if each operation spends one tenth of a second inside the chosen boundary on average, there are ten operations inside it on average.

In a different stable period, the same completion rate and half a second of average time correspond to fifty operations inside. Average time means the total time divided by the number of operations. Little's law is the relationship between these averages: work in flight equals throughput multiplied by time.

RateLimitly does not insert its minimum into that formula. A minimum and an average are different numbers. The tracker therefore does not calculate an exact queue size or read the database's remaining memory. Communication delays, slow computation, and connection waiting can all affect the chosen measurement.

Use latency feedback to decide when adding work looks unwise, not to claim perfect knowledge of spare capacity. A sudden failure can happen before enough new reports arrive. Independent resource bounds and operational monitoring still have jobs to do.

Diagram walkthrough
  1. Stable example: throughput × average time = average in flight.

    For a stable period and a consistent boundary, average work in flight equals throughput multiplied by average time.

  2. A longer average time retains more work at the same throughput.

    Derived examples: 100 × 0.1 = 10; 100 × 0.5 = 50. Longer time can retain more simultaneous work even with the same completion rate.

  3. The tracker minimum is not the average in this equation.

    RateLimitly does not substitute its minimum sample into Little’s law. Minimum and average are different statistics.

  4. Combine feedback with independent safeguards.

    The signal is evidence for admission, not an exact measurement of queue length, spare memory, utilization, or stability.

Content revision: 07c7dd00db17

Download review copy (27 lessons)
Check your understanding & apply it

Can you substitute the tracker’s recent minimum for the average time in Little’s law?

Show the explained answer

No. Little’s law relates compatible averages under appropriate conditions. The tracker’s minimum is a different signal.

Integration guidance

Calibrate and test thresholds against real workload behavior while keeping independent resource limits and monitoring.

Where the promise stops

The 10- and 50-operation examples describe separate stable periods, not an exact estimate during a growing transient queue.

Technical depth & sources

Concepts used here

Implementation detail

Use consistent measurement boundaries and compatible averaging intervals. Neither latency alone nor a minimum establishes cause, utilization, memory, or stability.

Original explanation inspired by Fred Hébert and operational references; no endorsement implied.

AI-generated narration: ElevenLabs Eleven v3, Daniel stock voice. Audio streams only when you start listening.

ElevenLabs

Why teams choose us

Happier customers
Consistent response times when it matters most.
Faster delivery
Ship features, not DIY limiters and brittle ops workarounds.
Clear controls
Per‑tenant and per‑route limits your stakeholders can reason about.

Engineering blog

View all posts →

Why you shouldn’t use Redis as a rate limiter: Part 1 of 2

A tour of the common Redis-based rate limiter implementations — and the correctness and performance traps each one hides.

Auto-Scaling Won’t Save You

The myth of infinite serverless scale — why adding machines doesn’t fix overload, and what to do instead.