Docs/API
Errors, and what each one charged
What does each status mean, and did it cost me anything?
Checked against the code on
Errors use standard HTTP status codes and the OpenAI error envelope. The column that matters most is the last one: some failures happen after the model has already run, and those are charged.
{ "error": { "message": "Not enough credits. This request needs about 12. Add credits at https://heyaskr.ai/wallet", "type": "insufficient_quota", "param": null, "code": null } }Before the model runs
None of these charge anything. The reservation, if one was made, is released.
| Status | Type | Message | What to do |
|---|---|---|---|
| 400 | invalid_request_error | messages is required and must be a non-empty array. | Send a messages array. |
| 400 | invalid_request_error | Unsupported message role: X | Use system, user or assistant. |
| 400 | invalid_request_error | No usable messages were provided. | Send at least one message with string content. |
| 400 | invalid_request_error | The conversation is too long. | Trim below the model's own context window (120,000 characters on a model with a small window). |
| 400 | invalid_request_error | That model has no published rate and cannot be billed. | Pick a model from the catalog. |
| 401 | authentication_error | Missing or invalid API key. | Check the header; make a new key if in doubt. |
| 402 | insufficient_quota | Not enough credits. This request needs about N. ... | Add credits. N is the worst case for this request, so a smaller max_tokens can get it through. |
| 404 | invalid_request_error | The model 'x' does not exist. | Ids are exact. Copy from the catalog. |
| 429 | rate_limit_error | This key's daily cap of N credits would be exceeded. ... | Raise the cap in the wallet or wait for the window to roll. |
| 429 | (plain body) | {"statusCode":429,"error":"Too Many Requests","message":"Rate limit exceeded, retry in 1 minute"} | 120 per minute per IP. This one is not in the OpenAI envelope. Back off and retry. |
| 400 | invalid_request_error | web_search is not available on this server. (code W206) | This server has no search. Drop web_search. param is web_search, code is W206. |
| 400 | invalid_request_error | This model cannot use web_search; see supportsTools on GET /v1/models. (code W207) | Pick a model with supportsTools: true, or drop web_search. code is W207. |
| 429 | rate_limit_error | Search the web is limited to a few turns a minute. Wait a moment and send again. (code W109) | 6 searching calls a minute per account, workspace and API together. Wait and retry. code is W109. |
| 400 | invalid_request_error | The model 'x' does not accept image input. | An inline picture on a model whose inputs on GET /v1/models lack image. Pick one that sees pictures. |
| 400 | invalid_request_error | The model 'x' does not accept file input. Upload the file with POST /v1/files and send its file_id instead; ... | An inline file, input_audio or video_url part on a model that takes no such input. Upload it and send file_id, and a PDF or sound goes as words. Files. |
| 400 | invalid_request_error | Unsupported content part: x. | Only text, image_url, file, input_audio and video_url parts are read. |
| 404 | invalid_request_error | No file with id 'x'. | The file_id names nothing on your account. Upload it again. |
| 413 | invalid_request_error | At most 6 pdf files per request. | Over the per-request count for that kind. Limits. |
| 413 | invalid_request_error | This request, with its files, is about N tokens; the model 'x' takes M. | Over the model's window. code is context_length_exceeded. Pick a bigger window, or send less. |
| 400 | invalid_request_error | That model cannot watch video. Pick one that does, or remove the video. | An upload this model cannot take. code carries the workspace code (W609 here; also W203, W605, W606, W607, W608, W611). Error codes. |
| 415, 413 or 400 | invalid_request_error | That file type is not accepted. ..., That file is too large. The limit is ..., That file could not be read. ... | An upload refused at POST /v1/files, with W604, W605 or W610 in code. Files. |
| 503 | api_error | Model access is temporarily unavailable. Nothing was charged. | Retry shortly. |
| 502 | api_error | The model gateway is unavailable. | Retry with backoff. |
After the model has run
| Status | Type | Message | Charged? | What to do |
|---|---|---|---|---|
| 429 | rate_limit_error | The upstream gateway is busy. Retry shortly. | No | Retry after a short delay. |
| 400 | invalid_request_error | The model's provider declined this request under its content rules. Asking for a real public figure is a common cause. Nothing was charged. Reword it, or use another model. Its code is content_policy_violation, as on OpenAI. | No | Reword the prompt or use another model. Retrying it unchanged is refused again. |
| 502 or 400 | api_error | The model gateway rejected the request. | No | Any other provider failure. 502 when the provider failed, 400 when it turned the request down for a reason other than its content rules. Check the prompt and the model. |
| 502 | api_error | The model returned no content. Nothing was charged. | No | A request that returns nothing is an error, not an empty answer. Retry or change the prompt. |
| 502 | api_error | The gateway sent an unreadable response. It was billed upstream, so the reserved credits were charged. | Yes, the reservation | Do not retry blindly. Check your activity first. |
| 504 | api_error | The model took longer than the gateway deadline. It was generated and billed upstream, so the reserved credits were charged. | Yes, the reservation | Lower max_tokens or pick a faster model. Do not retry blindly. |
With stream: true the refusal can also arrive after the stream has started, as an SSE event with the same code. The stream then closes without [DONE], and it is settled like any stream that ends early: on what it carried, which includes the prompt. That is why this message leaves out "Nothing was charged". Any other provider error mid-stream arrives the same way, with type api_error and code upstream_stream_error.
data: {"error":{"message":"The model's provider declined this request under its content rules. Asking for a real public figure is a common cause. Reword it, or use another model.","type":"invalid_request_error","param":null,"code":"content_policy_violation"}}A retry loop that treats every 5xx as free will double-spend on the two charged cases. Retry 503, 502 with unavailable or rejected, and 429. Log 504 and the unreadable 502 and look before retrying. Never retry content_policy_violation unchanged: the same request is refused again.
Backoff that behaves
import time, requests
RETRY = {429, 503}
def call(payload, tries=4):
for attempt in range(tries):
r = requests.post("https://heyaskr.ai/v1/chat/completions", json=payload,
headers={"Authorization": "Bearer askr_live_..."}, timeout=130)
if r.status_code == 200:
return r.json()
body = r.json().get("error", {})
charged = "reserved credits were charged" in body.get("message", "")
if r.status_code in RETRY or (r.status_code == 502 and not charged):
time.sleep(2 ** attempt)
continue
raise RuntimeError(f"{r.status_code}: {body.get('message')}")