Skip to main content

Keep your API fast—even at peak

Rate limiting and admission checks that help reduce overload and protect expensive work.

Reliable at peak

Set admission budgets for launch spikes and promos before excess work reaches your handlers.

Lower cloud costs

Fewer moving parts to run and monitor. Spend less on Redis and ops.

Dev‑first adoption

Add via SDK. Keep your stack. Clear limits your customers understand.

UNDERSTAND OVERLOAD · KEEP USEFUL WORK MOVING

Keep useful work moving.

For developers, architects, and managers: 27 narrated lessons on why services overload and how to protect them. No cloud experience required. Start with one request, or jump to any lesson.

Course map Lesson 25 of 27 Browse all lessons

Four chapters, one learning path. Read or listen in any order. ~41 minutes total · Times shown at 1×.

Keep enough capacity to finish work and recover.

  1. 01 A useful result starts with one request
  2. 02 Unfinished work grows when arrivals outrun completions
  3. 03 More waiting space does not make work finish faster
  4. 04 A timeout ends waiting, not necessarily the work
  5. 05 Retained work can exhaust a process’s memory
  6. 06 One failed copy can push its neighbors over capacity
  7. 07 Refuse new work while the application can still recover
  8. 08 Different compute platforms still have finite resources
  9. 09 A shared database failure can spread back into every app
  10. 10 Restore capacity before restoring full traffic
  11. 11 A circuit breaker stops repeating a failing call

Contain one customer's excess without punishing everyone.

  1. 12 One customer’s excess should not ruin another’s day
  2. 13 A rate limit checks one customer’s recent allowance
  3. 14 Copies and languages must share the same intended policy
  4. 15 Fairness helps contain damage without reserving every resource
  5. 16 A legitimate crowd can exceed shared capacity

Reduce work entering a shared dependency when it slows.

  1. 17 Measure the operation whose pressure you want to observe
  2. 18 A shared tracker observes; a guard decides; your app acts
  3. 19 Even the fastest recent observations can become slow
  4. 20 Separate local histories can reopen into the same burst
  5. 21 Replace optional expensive work before refusing essential work
  6. 22 A local breaker learns after calls have already piled up
  7. 23 Fast errors explain why the protections remain complementary
  8. 24 Latency feedback helps without measuring exact spare capacity

Choose, combine, and test the right protections.

  1. 25 Ask three different questions before expensive work You are here
  2. 26 Recovery attempts need limits too
  3. 27 Measure useful service, not how much work you accepted

Read the explanation or use the course map above. Audio is optional.

25 / 27

Ask three different questions before expensive work

Three checks before work

AllowRefuseLocal headroomTenant allowanceShared latency guardDo work · reportFallback or refusal
  1. Application-owned Local headroom?

    Memory and resource bounds

  2. RateLimitly rate limit Tenant within allowance?

    Fairness boundary

  3. RateLimitly latency guard Shared signal acceptable?

    Before dependency work

Conceptual checks, not three separate network calls. The application acts on refusal.

Read the numbered steps in order. The explanation below follows the narration.

Three different questions protect three different boundaries.

Combine local survival, tenant fairness, and dependency feedback without confusing denial with missing decisions.

0:00
Read transcript & diagram walkthrough

What you will learn

Combine local survival, tenant fairness, and dependency feedback without confusing denial with missing decisions.

Before expensive work, ask three different questions. Does this application have enough local headroom? Is this tenant within its allowance? Is the shared dependency's latency signal below this operation's threshold? The first question is application-owned local protection. The other two are RateLimitly's fairness and latency decisions.

If these required checks admit the work, perform the operation, measure its duration, and report it to the correct tracker. A policy refusal means a check answered no. Use a safe cheap fallback where the product permits one, or return an explicit refusal. Do not secretly start the refused operation anyway.

A transport failure is different: your application could not obtain a reliable answer from the decision service. A failure policy specifies what to do then. Allowing work despite that uncertainty favors immediate access; refusing favors protecting the dependency. Neither choice is universally correct. Choose and test it deliberately, with local bounds still active.

Keep the API key on the trusted application side, match identities and settings, and observe failures separately from ordinary policy refusals. The decision service also needs capacity and failure testing. A protection that gates work becomes part of the service you operate.

Diagram walkthrough
  1. Three different questions protect three different boundaries.

    Local resource safety, tenant fairness, and dependency pressure answer different questions. No one check establishes all three.

  2. Admitted work is measured; refused work does not start.

    Only start protected work after required checks allow it. A refusal must branch around the work, not merely add a log entry.

  3. No reliable answer is not a known refusal.

    An unavailable decision service is not an ordinary policy refusal. Choose and test failure behavior per operation while retaining independent local bounds.

  4. Operate the guardrails as part of the service.

    Keep identity, state scope, credentials and observed errors explicit. The decision service is itself a dependency whose failure behavior needs testing.

Content revision: 07c7dd00db17

Download review copy (27 lessons)
Check your understanding & apply it

A decision-service call failed. Is that the same as a policy refusal?

Show the explained answer

No. A refusal is a known decision. A transport failure is missing or unreliable decision evidence and needs an explicit failure policy.

Integration guidance

Local resource check → tenant/latency admission → branch → protected work → measured report. Keep credentials server-side and test refusal and transport-failure branches.

Where the promise stops

RateLimitly does not currently provide the application’s memory guard. A decision-service failure cannot safely be interpreted as automatic admission in every product.

Technical depth & sources

Concepts used here

Implementation detail

Choose fail-open or fail-closed per operation, observe admission overhead and failures, and keep independent bounds when decisions are unavailable.

Original explanation inspired by Fred Hébert and operational references; no endorsement implied.

AI-generated narration: ElevenLabs Eleven v3, Daniel stock voice. Audio streams only when you start listening.

ElevenLabs

Why teams choose us

Happier customers
Consistent response times when it matters most.
Faster delivery
Ship features, not DIY limiters and brittle ops workarounds.
Clear controls
Per‑tenant and per‑route limits your stakeholders can reason about.

Engineering blog

View all posts →

Why you shouldn’t use Redis as a rate limiter: Part 1 of 2

A tour of the common Redis-based rate limiter implementations — and the correctness and performance traps each one hides.

Auto-Scaling Won’t Save You

The myth of infinite serverless scale — why adding machines doesn’t fix overload, and what to do instead.