Docs/API

Limits and caps

How fast can I go, how big can a request be, and how do I stop a script overspending?

Checked against the code on

LimitValueWhere it bites
Requests per minute120 per IP address429 with no message body change. Back off.
Request body12 MB on chat completions, 100 MB on a file upload, 256 KB elsewhere413
Uploads per minute60 per IP address429
Files per request20 pictures, 6 PDFs, 10 text files, 6 documents, 4 audio files, 2 videos, across every message sent; per message 6, 3, 5, 3, 2 and 1413. Files
A file8 MB for a picture, 32 MB for a PDF or a document, 5 MB for text, 25 MB or 20 minutes for audio, 100 MB or 10 minutes for video413 at upload, code W605
ConversationLast 400 messages, up to the model's own context window (about 1 million tokens on the largest). A model with a small window: last 40 messages, 120,000 charactersOlder messages dropped; over the limit is 400
Outputmax_tokens up to 8,192, default 4,096Clamped, never rejected
Upstream deadline120 seconds504, charged, because the model ran
Active keys20 per account; 10 created per hour400 from the wallet
Daily cap per keyOptional, in credits, rolling 24 hours429 rate_limit_error before the call
Searching calls6 per minute per account, workspace and API together429 rate_limit_error, code W109, before the call
Search the web, per call3 searches, 2 pages read, 4 model callsThe model is told the limit is reached and answers with what it has. Search the web

There is no per-account request limit

The IP limit protects the service. The daily cap protects your balance, and it is the one to set. It is checked against the worst case a request could cost before the request runs, so a capped key cannot go over, even by one request.

Need more

Higher IP limits for a server that fans out on behalf of many users: email ask@heyaskr.ai with the address and the shape of the traffic.