> ## Documentation Index
> Fetch the complete documentation index at: https://docs.ycloud.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> Check YCloud API quotas, read rate limit headers, and retry throttled requests safely.

YCloud limits API requests within a time window. Limits apply to an account, a
sender, or a group of endpoints. Requests that exceed a limit return HTTP
`429 Too Many Requests` with the standard [error response](/en/api-reference/guides/api-fundamentals/handle-errors).

## Messaging API limits

`rps` means requests per second. A sender is a WhatsApp business phone number.

| Endpoint | Rate limit | Scope |
| - | - | - |
| `POST /v2/emails` | 200 rps | Per account |
| `POST /v2/sms` | 200 rps | Per account |
| `POST /v2/voices` | 200 rps | Per account |
| `POST /v2/whatsapp/messages` | 200 rps | Per sender |
| `POST /v2/whatsapp/messages/sendDirectly` | 80 rps by default; 1,000 rps after an eligible automatic throughput upgrade | Per sender |

Most endpoints have a limit of 200 rps per account. WhatsApp sending endpoints use
the sender limits shown above. See Meta's [throughput documentation](https://developers.facebook.com/docs/whatsapp/cloud-api/overview#throughput)
for automatic upgrades.

### Queued and direct WhatsApp sends

YCloud measures the two WhatsApp sending endpoints separately. The queued
`/v2/whatsapp/messages` endpoint accepts up to 200 rps per sender, while YCloud
submits queued messages to Meta at 60 rps. Acceptance into the queue does not
mean that Meta has accepted or delivered the message.

Avoid sending at a high rate through both endpoints with the same sender at the
same time. Queued and direct sends still use the same WhatsApp business phone
number, so their combined traffic can reach Meta's limit and cause failures.

## Management API limits

Most APIs without a separate documented policy are management APIs. These
endpoints share an account quota: **200 requests per second and 10,000 requests
per hour**. Both limits apply; you cannot sustain 200 rps for an entire hour.
Requests to one management endpoint consume quota available to the others.

The shared policy includes:

* `/v2/balance`
* `/v2/webhookEndpoints/*`
* `/v2/whatsapp/businessAccounts/*`
* `/v2/whatsapp/phoneNumbers/*`
* `/v2/whatsapp/templates/*`
* `/v2/whatsapp/messages/{id}`
* Other APIs without a separately documented policy

## Read rate limit headers

YCloud includes rate limit information in most API responses. Read the headers
when present instead of assuming that every endpoint has the same quota.

| Header | Meaning |
| - | - |
| `Retry-After` | Seconds to wait before retrying or sending another request. |
| `RateLimit-Limit` | Maximum quota units for the account or sender in the reported time window. |
| `RateLimit-Policy` | Informational quota policies and their time windows. For example, `100;w=60` describes 100 quota units per 60 seconds. |
| `RateLimit-Remaining` | Quota units still available for the reported limit. |
| `RateLimit-Reset` | Seconds until the reported quota resets. |

<Note>
  The `RateLimit-*` headers are beta and may change. YCloud's documented header
  format follows the IETF draft-06 rate limit specification. Treat policy
  parameters as informational and keep your client tolerant of additional
  parameters.
</Note>

### Example: shared hourly quota exhausted

```http theme={"theme":{"light":"github-light","dark":"github-dark"}}
HTTP/2 429
Content-Type: application/json
Retry-After: 1800
RateLimit-Limit: 10000
RateLimit-Policy: 200;w=1;burst=200;algorithm=token_bucket;level=account;scope=management_api, 10000;w=3600;algorithm=fixed_window;level=account;scope=management_api
RateLimit-Remaining: 0
RateLimit-Reset: 1800
```

The account has exhausted its shared **10,000 requests per hour** quota. Wait at
least 1,800 seconds before sending another request against that quota. Switching
to a different management endpoint does not provide a fresh quota.

## Handle a `429` response

1. Pause requests that share the exhausted account or sender quota.
2. Respect `Retry-After` when present. Do not retry before that delay expires.
3. Reduce concurrency and add exponential backoff with jitter.
4. Bound retries by an attempt count or your application's deadline.
5. Log the endpoint, HTTP method, error `code`, `requestId`, and request time.
   Exclude API keys and personal data.

If a successful response includes `Retry-After`, delay subsequent requests; do
not resend the request that already succeeded. Monitor `RateLimit-Remaining`
and `RateLimit-Reset` to slow down before the quota is exhausted.

### Retry with backoff and jitter

This JavaScript example treats `Retry-After` as a minimum wait. The backoff cap
limits the application's jitter delay; it does not shorten a server-requested
wait. The attempt and delay settings are application choices, not API limits.

```javascript theme={"theme":{"light":"github-light","dark":"github-dark"}}
async function requestWithBackoff(url, options) {
  const maxAttempts = 5;
  const baseDelayMs = 500;
  const maxBackoffMs = 30_000;

  for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
    const response = await fetch(url, options);
    if (response.status !== 429) return response;
    if (attempt === maxAttempts - 1) {
      throw new Error("YCloud API rate limit persisted after retries");
    }

    const header = response.headers.get("Retry-After");
    const seconds = header === null ? NaN : Number(header);
    const serverDelayMs = Number.isFinite(seconds) && seconds >= 0
      ? seconds * 1000
      : 0;
    const backoffCap = Math.min(maxBackoffMs, baseDelayMs * 2 ** attempt);
    const delayMs = Math.max(serverDelayMs, Math.random() * backoffCap);
    await response.body?.cancel();
    await new Promise((resolve) => setTimeout(resolve, delayMs));
  }
}
```

Apply this example only when the operation is safe to retry. For long waits, use
a scheduled queue so a worker does not need to stay active.

## Control concurrency and retries

Use a bounded queue and a shared worker limit. Avoid independent retry loops
that each assume the full account or sender quota is available. Restore traffic
gradually after throttling ends.

A repeated `POST` can create another resource or send another message. Check the
endpoint's retry guidance and your stored result first. Where supported, keep a
stable `externalId`, but do not treat it as a universal idempotency key. Store the
YCloud response ID and reconcile ambiguous outcomes before sending again.

Monitor request volume, `429` responses, latency, retry counts, and queue age by
endpoint and sender. See [WhatsApp Messages API best practices](/en/api-reference/guides/whatsapp-platform/whatsapp-messages-api-best-practices)
for queueing, reconciliation, and duplicate prevention.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.