Skip to main content
YCloud limits API requests within a time window. Limits apply to an account, a sender, or a group of endpoints. Requests that exceed a limit return HTTP 429 Too Many Requests with the standard error response.

Messaging API limits

rps means requests per second. A sender is a WhatsApp business phone number. Most endpoints have a limit of 200 rps per account. WhatsApp sending endpoints use the sender limits shown above. See Meta’s throughput documentation for automatic upgrades.

Queued and direct WhatsApp sends

YCloud measures the two WhatsApp sending endpoints separately. The queued /v2/whatsapp/messages endpoint accepts up to 200 rps per sender, while YCloud submits queued messages to Meta at 60 rps. Acceptance into the queue does not mean that Meta has accepted or delivered the message. Avoid sending at a high rate through both endpoints with the same sender at the same time. Queued and direct sends still use the same WhatsApp business phone number, so their combined traffic can reach Meta’s limit and cause failures.

Management API limits

Most APIs without a separate documented policy are management APIs. These endpoints share an account quota: 200 requests per second and 10,000 requests per hour. Both limits apply; you cannot sustain 200 rps for an entire hour. Requests to one management endpoint consume quota available to the others. The shared policy includes:
  • /v2/balance
  • /v2/webhookEndpoints/*
  • /v2/whatsapp/businessAccounts/*
  • /v2/whatsapp/phoneNumbers/*
  • /v2/whatsapp/templates/*
  • /v2/whatsapp/messages/{id}
  • Other APIs without a separately documented policy

Read rate limit headers

YCloud includes rate limit information in most API responses. Read the headers when present instead of assuming that every endpoint has the same quota.
The RateLimit-* headers are beta and may change. YCloud’s documented header format follows the IETF draft-06 rate limit specification. Treat policy parameters as informational and keep your client tolerant of additional parameters.

Example: shared hourly quota exhausted

The account has exhausted its shared 10,000 requests per hour quota. Wait at least 1,800 seconds before sending another request against that quota. Switching to a different management endpoint does not provide a fresh quota.

Handle a 429 response

  1. Pause requests that share the exhausted account or sender quota.
  2. Respect Retry-After when present. Do not retry before that delay expires.
  3. Reduce concurrency and add exponential backoff with jitter.
  4. Bound retries by an attempt count or your application’s deadline.
  5. Log the endpoint, HTTP method, error code, requestId, and request time. Exclude API keys and personal data.
If a successful response includes Retry-After, delay subsequent requests; do not resend the request that already succeeded. Monitor RateLimit-Remaining and RateLimit-Reset to slow down before the quota is exhausted.

Retry with backoff and jitter

This JavaScript example treats Retry-After as a minimum wait. The backoff cap limits the application’s jitter delay; it does not shorten a server-requested wait. The attempt and delay settings are application choices, not API limits.
Apply this example only when the operation is safe to retry. For long waits, use a scheduled queue so a worker does not need to stay active.

Control concurrency and retries

Use a bounded queue and a shared worker limit. Avoid independent retry loops that each assume the full account or sender quota is available. Restore traffic gradually after throttling ends. A repeated POST can create another resource or send another message. Check the endpoint’s retry guidance and your stored result first. Where supported, keep a stable externalId, but do not treat it as a universal idempotency key. Store the YCloud response ID and reconcile ambiguous outcomes before sending again. Monitor request volume, 429 responses, latency, retry counts, and queue age by endpoint and sender. See WhatsApp Messages API best practices for queueing, reconciliation, and duplicate prevention.