Skip to main content

Rate Limits

Rate limits are enforced per API key based on your subscription tier:

Rate Limit Headers

Every response includes rate limit headers:

Handling 429 Errors

When rate limited, you’ll receive a 429 Too Many Requests response:

Automatic Retries

All official SDKs automatically retry on 429 errors with exponential backoff:
The retry delay follows: min(500ms × 2^attempt, 5000ms)

Best Practices

  1. Use exponential backoff — the SDKs handle this automatically
  2. Monitor remaining quota — check X-RateLimit-Remaining headers
  3. Use streaming — streaming requests count as one request regardless of response length
  4. Upgrade your tier — if you consistently hit limits, consider upgrading
  5. Distribute load — use multiple API keys for different services