Rate limits
The API uses token-bucket rate
limiting. Each limit allows a short burst and a sustained request rate; requests
beyond an applicable limit receive 429 Too Many Requests.
How Token Buckets Work
Each limit is a bucket with two numbers:
- Capacity — the maximum burst. The bucket starts full; each admitted operation consumes one token from each applicable operation bucket.
- Refill / sec — the sustained rate at which tokens are replenished.
If at least one token is available, the request is allowed and a token is consumed. If the bucket is empty, the request is rejected with a 429 that tells you exactly how long to wait via retry_after. An idle bucket refills back to full, so you regain your full burst allowance after a quiet period.
Example: a bucket with capacity 8 and refill 4/s lets you fire 8 requests back-to-back, then accepts roughly 4 more per second sustained. If you drain it, waiting one second restores 4 tokens.
Limit Levels
Every admitted operation is checked against each applicable operation bucket. Authentication failures do not consume a token. An idempotency replay returns the stored response without re-executing or consuming a token.
Buckets are normally per command type. At the per-command and per-room levels, command types normally have their own bucket. The exception is the shared seek group: playback/seek, recording/mask, and recording/unmask draw from one 20 / 20 bucket. A burst of one of those commands can drain the budget available to the others; other command types use independent buckets.
The 429 Response
When a hard limit is exceeded, the gateway returns 429 Too Many Requests with a retry_after field telling you how many seconds to wait before the bucket has a token again:
Branch on code (the stable enum, always "rate_limited" for a 429), not the human-readable error string, which may change. retry_after is an integer number of seconds; the Retry-After header mirrors that value. Wait at least that long, then retry. Include the traceparent response header when contacting support about a request.
Per-Command Limits
The table below lists each command’s per-call bucket as
capacity / refill-per-second. The three commands in the shared seek group
draw from the same bucket rather than receiving a bucket apiece.
Room-member volume handlers do not use per-command-in-call buckets. Room-member add and delete endpoints use only the per-room limits below.
Per-Room Limits
Room commands are additionally checked against a per-room bucket that aggregates across all members of the room.
Per-App Limits (App-Global Commands)
Creating a room and listing rooms or sessions are limited per application, aggregated across all requests. Each limit is shown as capacity / refill-per-second.
Resource Limits (Not a Rate Limit)
This is a count cap, not a rate cap. Everything above bounds how fast you may call; this bounds how many rooms exist at once. The two are independent: staying under the creation rate limit does not exempt you from the room-count cap, and vice versa.
Each application has a cap on the number of rooms it may have concurrently live at one time. The cap counts every room your app currently owns, including rooms that are still being torn down — a room frees its slot only once it is deleted and its teardown has finalized.
The default room cap is 100 concurrently live rooms. Contact support for
higher limit.
When creating a room would exceed the cap, POST /api/v1/rooms is rejected with 403 Forbidden — not a 429. Because this is not a rate limit, there is no retry_after: the condition clears when you delete a room and free a slot, not after a fixed delay.
Room member cap
A second count cap bounds how many members may be in a single room at once. It is independent of the room-count cap above and of every rate limit. The default is 250 members per room. Contact support for higher limit.
When a join (POST /api/v1/rooms/{room_id}/members) would exceed the cap, it is rejected with 409 Conflict and code: room_full — not a 429 and not the room-count 403. The member is not seated. As with the room-count cap this is not a rate limit, so there is no retry_after: the condition clears when a member leaves and frees a slot, not after a fixed delay.
The cap applies to members added through the join endpoint.
Best Practices
- Honor
retry_after. On a429, wait the number of seconds inretry_afterbefore retrying. Requests above the refill rate receive429responses. - Wait for webhooks. For asynchronous operations, use the lifecycle webhook to detect completion instead of polling.
- Batch playback files. Send multiple URLs in one
playback/playcommand withtype: files. - Back off on repeated 429s. If a retry receives another
429, wait for its newretry_aftervalue before retrying again.