Rate limits

Understand request limits and handle traffic safely.
View as Markdown

The API uses token-bucket rate limiting. Each limit allows a short burst and a sustained request rate; requests beyond an applicable limit receive 429 Too Many Requests.

How Token Buckets Work

Each limit is a bucket with two numbers:

  • Capacity — the maximum burst. The bucket starts full; each admitted operation consumes one token from each applicable operation bucket.
  • Refill / sec — the sustained rate at which tokens are replenished.

If at least one token is available, the request is allowed and a token is consumed. If the bucket is empty, the request is rejected with a 429 that tells you exactly how long to wait via retry_after. An idle bucket refills back to full, so you regain your full burst allowance after a quiet period.

Example: a bucket with capacity 8 and refill 4/s lets you fire 8 requests back-to-back, then accepts roughly 4 more per second sustained. If you drain it, waiting one second restores 4 tokens.

Limit Levels

Every admitted operation is checked against each applicable operation bucket. Authentication failures do not consume a token. An idempotency replay returns the stored response without re-executing or consuming a token.

LevelApplies toEnforced?
Per-command-in-callEach command type on a call, normally with its own bucket; playback/seek, recording/mask, and recording/unmask share the seek bucket.Yes — 429
Per-roomRoom commands, aggregated across all members of a room, per command type.Yes — 429
Per-appApp-global commands not tied to a session — room/conference creation and listings — aggregated per application.Yes — 429

Buckets are normally per command type. At the per-command and per-room levels, command types normally have their own bucket. The exception is the shared seek group: playback/seek, recording/mask, and recording/unmask draw from one 20 / 20 bucket. A burst of one of those commands can drain the budget available to the others; other command types use independent buckets.

The 429 Response

When a hard limit is exceeded, the gateway returns 429 Too Many Requests with a retry_after field telling you how many seconds to wait before the bucket has a token again:

1HTTP/1.1 429 Too Many Requests
2Retry-After: 1
3
4{
5 "success": false,
6 "error": "Command rate limit exceeded",
7 "code": "rate_limited",
8 "retry_after": 1
9}

Branch on code (the stable enum, always "rate_limited" for a 429), not the human-readable error string, which may change. retry_after is an integer number of seconds; the Retry-After header mirrors that value. Wait at least that long, then retry. Include the traceparent response header when contacting support about a request.

Per-Command Limits

The table below lists each command’s per-call bucket as capacity / refill-per-second. The three commands in the shared seek group draw from the same bucket rather than receiving a bucket apiece.

CommandCapacity / Refill (per sec)Notes
playback/seek20 / 20Shares the seek bucket with recording/mask and recording/unmask.
playback/stop20 / 20
playback/pause20 / 20
playback/resume20 / 20
playback/restart20 / 20
recording/stop20 / 20
recording/mask20 / 20Shares the seek bucket with playback/seek and recording/unmask.
recording/unmask20 / 20Shares the seek bucket with playback/seek and recording/mask.
playback/play8 / 4Plays files, a live stream, or silence.
play_and_get_digits (POST + DELETE)4 / 0.8The prompt-and-collect POST and cancel DELETE each have an independent bucket with the same capacity and refill rate.
recording/start4 / 1Starts a recording.
room/mute (conference.mute)15 / 10
room/unmute (conference.unmute)15 / 10

Room-member volume handlers do not use per-command-in-call buckets. Room-member add and delete endpoints use only the per-room limits below.

Per-Room Limits

Room commands are additionally checked against a per-room bucket that aggregates across all members of the room.

Room commandCapacity / Refill (per sec)Scope
POST /rooms/{room_id}/members100 / 30per-room
DELETE /rooms/{room_id}/members/{uuid}100 / 30per-room
POST /rooms/{room_id}/playback/pause15 / 10per-room
POST /rooms/{room_id}/playback/seek15 / 10per-room
POST /rooms/{room_id}/playback/volume15 / 10per-room
POST /rooms/{room_id}/playback/stop15 / 10per-room
POST /rooms/{room_id}/playback/play8 / 3per-room
POST /rooms/{room_id}/playback/record8 / 3per-room; start and stop share this bucket
POST /rooms/{room_id}/playback/record/stop8 / 3per-room; start and stop share this bucket

Per-App Limits (App-Global Commands)

Creating a room and listing rooms or sessions are limited per application, aggregated across all requests. Each limit is shown as capacity / refill-per-second.

CommandHard limitNotes
POST /rooms (create)10 / 2per-app — room/conference creation
GET /sessions (list)40 / 20per-app
GET /rooms (list)40 / 20per-app

Resource Limits (Not a Rate Limit)

This is a count cap, not a rate cap. Everything above bounds how fast you may call; this bounds how many rooms exist at once. The two are independent: staying under the creation rate limit does not exempt you from the room-count cap, and vice versa.

Each application has a cap on the number of rooms it may have concurrently live at one time. The cap counts every room your app currently owns, including rooms that are still being torn down — a room frees its slot only once it is deleted and its teardown has finalized.

The default room cap is 100 concurrently live rooms. Contact support for higher limit.

When creating a room would exceed the cap, POST /api/v1/rooms is rejected with 403 Forbiddennot a 429. Because this is not a rate limit, there is no retry_after: the condition clears when you delete a room and free a slot, not after a fixed delay.

1HTTP/1.1 403 Forbidden
2
3{
4 "success": false,
5 "error": "Maximum number of rooms reached for this application",
6 "code": "forbidden"
7}

Room member cap

A second count cap bounds how many members may be in a single room at once. It is independent of the room-count cap above and of every rate limit. The default is 250 members per room. Contact support for higher limit.

When a join (POST /api/v1/rooms/{room_id}/members) would exceed the cap, it is rejected with 409 Conflict and code: room_fullnot a 429 and not the room-count 403. The member is not seated. As with the room-count cap this is not a rate limit, so there is no retry_after: the condition clears when a member leaves and frees a slot, not after a fixed delay.

1HTTP/1.1 409 Conflict
2
3{
4 "success": false,
5 "error": "Room is full",
6 "code": "room_full"
7}

The cap applies to members added through the join endpoint.

Best Practices

  • Honor retry_after. On a 429, wait the number of seconds in retry_after before retrying. Requests above the refill rate receive 429 responses.
  • Wait for webhooks. For asynchronous operations, use the lifecycle webhook to detect completion instead of polling.
  • Batch playback files. Send multiple URLs in one playback/play command with type: files.
  • Back off on repeated 429s. If a retry receives another 429, wait for its new retry_after value before retrying again.