429 Too Many Requests with the standard error response.
Messaging API limits
rps means requests per second. A sender is a WhatsApp business phone number.
Most endpoints have a limit of 200 rps per account. WhatsApp sending endpoints use
the sender limits shown above. See Meta’s throughput documentation
for automatic upgrades.
Queued and direct WhatsApp sends
YCloud measures the two WhatsApp sending endpoints separately. The queued/v2/whatsapp/messages endpoint accepts up to 200 rps per sender, while YCloud
submits queued messages to Meta at 60 rps. Acceptance into the queue does not
mean that Meta has accepted or delivered the message.
Avoid sending at a high rate through both endpoints with the same sender at the
same time. Queued and direct sends still use the same WhatsApp business phone
number, so their combined traffic can reach Meta’s limit and cause failures.
Management API limits
Most APIs without a separate documented policy are management APIs. These endpoints share an account quota: 200 requests per second and 10,000 requests per hour. Both limits apply; you cannot sustain 200 rps for an entire hour. Requests to one management endpoint consume quota available to the others. The shared policy includes:/v2/balance/v2/webhookEndpoints/*/v2/whatsapp/businessAccounts/*/v2/whatsapp/phoneNumbers/*/v2/whatsapp/templates/*/v2/whatsapp/messages/{id}- Other APIs without a separately documented policy
Read rate limit headers
YCloud includes rate limit information in most API responses. Read the headers when present instead of assuming that every endpoint has the same quota.The
RateLimit-* headers are beta and may change. YCloud’s documented header
format follows the IETF draft-06 rate limit specification. Treat policy
parameters as informational and keep your client tolerant of additional
parameters.Example: shared hourly quota exhausted
Handle a 429 response
- Pause requests that share the exhausted account or sender quota.
- Respect
Retry-Afterwhen present. Do not retry before that delay expires. - Reduce concurrency and add exponential backoff with jitter.
- Bound retries by an attempt count or your application’s deadline.
- Log the endpoint, HTTP method, error
code,requestId, and request time. Exclude API keys and personal data.
Retry-After, delay subsequent requests; do
not resend the request that already succeeded. Monitor RateLimit-Remaining
and RateLimit-Reset to slow down before the quota is exhausted.
Retry with backoff and jitter
This JavaScript example treatsRetry-After as a minimum wait. The backoff cap
limits the application’s jitter delay; it does not shorten a server-requested
wait. The attempt and delay settings are application choices, not API limits.
Control concurrency and retries
Use a bounded queue and a shared worker limit. Avoid independent retry loops that each assume the full account or sender quota is available. Restore traffic gradually after throttling ends. A repeatedPOST can create another resource or send another message. Check the
endpoint’s retry guidance and your stored result first. Where supported, keep a
stable externalId, but do not treat it as a universal idempotency key. Store the
YCloud response ID and reconcile ambiguous outcomes before sending again.
Monitor request volume, 429 responses, latency, retry counts, and queue age by
endpoint and sender. See WhatsApp Messages API best practices
for queueing, reconciliation, and duplicate prevention.
