Skip to main content

Keep your API fast—even at peak

Rate limiting and admission checks that help reduce overload and protect expensive work.

Reliable at peak

Set admission budgets for launch spikes and promos before excess work reaches your handlers.

Lower cloud costs

Fewer moving parts to run and monitor. Spend less on Redis and ops.

Dev‑first adoption

Add via SDK. Keep your stack. Clear limits your customers understand.

UNDERSTAND OVERLOAD · KEEP USEFUL WORK MOVING

Keep useful work moving.

For developers, architects, and managers: 27 narrated lessons on why services overload and how to protect them. No cloud experience required. Start with one request, or jump to any lesson.

Course map Lesson 14 of 27 Browse all lessons

Four chapters, one learning path. Read or listen in any order. ~41 minutes total · Times shown at 1×.

Keep enough capacity to finish work and recover.

  1. 01 A useful result starts with one request
  2. 02 Unfinished work grows when arrivals outrun completions
  3. 03 More waiting space does not make work finish faster
  4. 04 A timeout ends waiting, not necessarily the work
  5. 05 Retained work can exhaust a process’s memory
  6. 06 One failed copy can push its neighbors over capacity
  7. 07 Refuse new work while the application can still recover
  8. 08 Different compute platforms still have finite resources
  9. 09 A shared database failure can spread back into every app
  10. 10 Restore capacity before restoring full traffic
  11. 11 A circuit breaker stops repeating a failing call

Contain one customer's excess without punishing everyone.

  1. 12 One customer’s excess should not ruin another’s day
  2. 13 A rate limit checks one customer’s recent allowance
  3. 14 Copies and languages must share the same intended policy You are here
  4. 15 Fairness helps contain damage without reserving every resource
  5. 16 A legitimate crowd can exceed shared capacity

Reduce work entering a shared dependency when it slows.

  1. 17 Measure the operation whose pressure you want to observe
  2. 18 A shared tracker observes; a guard decides; your app acts
  3. 19 Even the fastest recent observations can become slow
  4. 20 Separate local histories can reopen into the same burst
  5. 21 Replace optional expensive work before refusing essential work
  6. 22 A local breaker learns after calls have already piled up
  7. 23 Fast errors explain why the protections remain complementary
  8. 24 Latency feedback helps without measuring exact spare capacity

Choose, combine, and test the right protections.

  1. 25 Ask three different questions before expensive work
  2. 26 Recovery attempts need limits too
  3. 27 Measure useful service, not how much work you accepted

Read the explanation or use the course map above. Audio is optional.

14 / 27

Copies and languages must share the same intended policy

Shared policy across languages

same API keysame tenant bucketsame window andallowancePython serviceRateLimitlyJavaScript serviceAnother applicationcopyOne intended tenantpolicy
  1. Python app One local allowance

    Same tenant · O

  2. JavaScript app Another local allowance

    Same tenant · O

  3. Effective policy Accidental multiplication

    More copies, more allowance

Illustrative mechanism. Highlight follows the explanation; not a benchmark.

Read the numbered steps in order. The explanation below follows the narration.

Separate local allowances can multiply one customer’s share.

Use the same API key, tenant bucket name, window, and allowance across every client.

0:00
Read transcript & diagram walkthrough

What you will learn

Use the same API key, tenant bucket name, window, and allowance across every client.

One customer may reach several application copies, written in different programming languages. If each copy gives Orchid its own full allowance, adding copies also multiplies Orchid's effective allowance. The policy has changed without anyone intending it.

RateLimitly is a service that applications ask for rate and latency decisions. A client library is code that lets your program make those calls. The API key defines which RateLimitly servers those libraries use. Keep it on the trusted application side.

A bucket is the named usage record for a rate-limit policy. Use the same API key, the same tenant-derived bucket name, and the same time window and allowance across the client libraries. The libraries then apply the intended policy across your application copies instead of giving each copy its own allowance.

For example, Orchid's report requests should use the same bucket whether they come through Python or JavaScript. Adding an application copy or changing languages should not grant Orchid another share. Use distinct tenant-derived names for different customers so one customer's use is accounted for separately.

Diagram walkthrough
  1. Separate local allowances can multiply one customer’s share.

    Independent full allowances per application copy multiply a tenant’s effective budget when the fleet grows.

  2. The API key defines which RateLimitly servers the libraries use.

    The API key defines which RateLimitly servers the libraries use. Keep the key in trusted application code, not the public browser.

  3. Use the same API key, tenant bucket name, window, and allowance.

    Use the same API key, tenant-derived bucket name, time window, and allowance across the libraries. Each application copy applies the intended policy.

  4. Changing languages does not give Orchid another share.

    Orchid uses the same named policy through Python and JavaScript. Different customers use distinct tenant-derived bucket names, keeping their usage separate.

Content revision: 07c7dd00db17

Download review copy (27 lessons)
Check your understanding & apply it

Orchid's report moves from Python to JavaScript. Should that give Orchid another allowance?

Show the explained answer

No. Use the same API key and the same tenant-derived bucket name, window, and allowance so both clients apply the same intended policy.

Integration guidance

Configure every client library with the same API key, tenant-derived bucket name, time window, and allowance. Keep the key on the trusted application side.

Where the promise stops

This example covers one tenant's named policy under the same API key. It does not reserve memory or execution slots.

Technical depth & sources

Concepts used here

Implementation detail

Bucket identity is derived from its logical name, time window, and rate limit. Changing those settings defines a different policy; each request's work weight does not change the bucket identity.

Original explanation inspired by Fred Hébert and operational references; no endorsement implied.

AI-generated narration: ElevenLabs Eleven v3, Daniel stock voice. Audio streams only when you start listening.

ElevenLabs

Why teams choose us

Happier customers
Consistent response times when it matters most.
Faster delivery
Ship features, not DIY limiters and brittle ops workarounds.
Clear controls
Per‑tenant and per‑route limits your stakeholders can reason about.

Engineering blog

View all posts →

Why you shouldn’t use Redis as a rate limiter: Part 1 of 2

A tour of the common Redis-based rate limiter implementations — and the correctness and performance traps each one hides.

Auto-Scaling Won’t Save You

The myth of infinite serverless scale — why adding machines doesn’t fix overload, and what to do instead.