Docs/API
Limits and caps
How fast can I go, how big can a request be, and how do I stop a script overspending?
Checked against the code on
| Limit | Value | Where it bites |
|---|---|---|
| Requests per minute | 120 per IP address | 429 with no message body change. Back off. |
| Request body | 12 MB on chat completions, 100 MB on a file upload, 256 KB elsewhere | 413 |
| Uploads per minute | 60 per IP address | 429 |
| Files per request | 20 pictures, 6 PDFs, 10 text files, 6 documents, 4 audio files, 2 videos, across every message sent; per message 6, 3, 5, 3, 2 and 1 | 413. Files |
| A file | 8 MB for a picture, 32 MB for a PDF or a document, 5 MB for text, 25 MB or 20 minutes for audio, 100 MB or 10 minutes for video | 413 at upload, code W605 |
| Conversation | Last 400 messages, up to the model's own context window (about 1 million tokens on the largest). A model with a small window: last 40 messages, 120,000 characters | Older messages dropped; over the limit is 400 |
| Output | max_tokens up to 8,192, default 4,096 | Clamped, never rejected |
| Upstream deadline | 120 seconds | 504, charged, because the model ran |
| Active keys | 20 per account; 10 created per hour | 400 from the wallet |
| Daily cap per key | Optional, in credits, rolling 24 hours | 429 rate_limit_error before the call |
| Searching calls | 6 per minute per account, workspace and API together | 429 rate_limit_error, code W109, before the call |
| Search the web, per call | 3 searches, 2 pages read, 4 model calls | The model is told the limit is reached and answers with what it has. Search the web |
There is no per-account request limit
The IP limit protects the service. The daily cap protects your balance, and it is the one to set. It is checked against the worst case a request could cost before the request runs, so a capped key cannot go over, even by one request.
Need more
Higher IP limits for a server that fans out on behalf of many users: email ask@heyaskr.ai with the address and the shape of the traffic.