Docs/API
Chat completions
Every field we read, every limit we apply, and what comes back.
Checked against the code on
POST https://heyaskr.ai/v1/chat/completionsCreate a model response for a conversation. The request is the OpenAI shape. This page is precise about which fields we read, because the rest are dropped without a warning.
Request fields we read
| Field | Type | What it does |
|---|---|---|
| model | string | A chat model id from the catalog. Exact match; there are no suffixes or aliases. |
| messages | array | The conversation. Each item has role (system, user or assistant) and content: a string, or an array of parts for pictures and files. Pictures and files. |
| max_tokens | integer | Output ceiling, thinking included. Default 4,096, maximum 8,192. Larger values are clamped, not rejected. |
| stream | boolean | true for server-sent events. Streaming. |
| web_search | boolean | askr's own field. true lets the model search the web and read pages during the call; the sources come back as annotations. Search the web. |
Everything else is dropped silently: temperature, top_p, tools, tool_choice, response_format, stop, n, seed, plugins. A request that sends them succeeds; they just have no effect. Tool calling and structured output are on the roadmap. Web search is not a tool you define: it is the web_search field above.
Messages
- Roles other than
system,userandassistantreturn400 Unsupported message role. contentis a string, or an array of parts:text,image_url,file,input_audioandvideo_url, each inline or byfile_id. Any other part type returns400 Unsupported content part. A message that is only parts is sent as it is. Pictures and files.- Only the last 40 messages are sent to the model. Trim on your side if you want control over what is kept.
- A conversation may fill the model's own context window, less room for the reply: about 1 million tokens on the largest models, such as Kimi K3, at most 3.2 million characters and the last 400 messages. A model with a small window keeps 120,000 characters and 40 messages. Over the limit returns
400 The conversation is too long. - No usable messages returns
400 No usable messages were provided.
Example
curl https://heyaskr.ai/v1/chat/completions \
-H "Authorization: Bearer $ASKR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5",
"max_tokens": 300,
"messages": [
{"role": "system", "content": "Answer in one sentence."},
{"role": "user", "content": "Why is the sky blue?"}
]
}'Response
{
"id": "chatcmpl-9c2b...",
"object": "chat.completion",
"created": 1789516800,
"model": "claude-sonnet-5",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Sunlight scatters off air molecules, and blue scatters most." },
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 21,
"completion_tokens": 14,
"total_tokens": 35,
"completion_tokens_details": { "reasoning_tokens": 0 }
},
"askr": { "credits_charged": 0.192 }
}| Field | Meaning |
|---|---|
| choices[0].message.content | The reply. Always one choice. |
| choices[0].message.images | Present only when the model answered with a picture. Images from chat. |
| choices[0].message.annotations | Present only on a web_search: true call: one url_citation per source. Search the web. |
| usage.completion_tokens_details.reasoning_tokens | On a thinking model, output tokens you never see. They are billed like any other output token. |
| askr.credits_charged | What the call cost, in credits, to four decimal places. Additive: clients that do not know it ignore it. |
How the cost is worked out
Before the call, credits are reserved for the worst case: your estimated input tokens at the model's input rate, plus max_tokens at its output rate. When the model finishes, the reservation is settled to the real cost, using the gateway's own figure for the call. Nothing else is added: there is no platform fee in the first version. Holds and settlement.
Pictures and files in a message
A user message's content can be an array of parts, in OpenAI's shape. Five kinds are read. A part carries its bytes inline, or names an upload by file_id from POST /v1/files; an upload is the way to send anything big, or anything more than once.
| Part | Inline | By upload | Inline limit |
|---|---|---|---|
| text | text | ||
| image_url | image_url.url: a PNG, JPEG, WebP or GIF data URL | image_url.file_id | 8 MB decoded; 6 a message, 20 a request |
| file | file.file_data: a PDF as a data URL, or an https URL the provider fetches. file.filename is optional. | file.file_id | 32 MB decoded |
| input_audio | input_audio.data as base64, with input_audio.format wav or mp3 | input_audio.file_id | 25 MB decoded |
| video_url | video_url.url: an MP4, MOV or WebM data URL, or an https URL | video_url.file_id | 25 MB decoded |
The request body on this endpoint is 12 MB, so an inline part over about 9 MB does not fit: upload it. An inline part goes to the model as it is, so it works only on a model whose inputs on GET /v1/models include that kind. A part by file_id is routed per model: the file itself where the model takes the kind, its words where it does not, so a PDF by file_id works on every model. Files has the rules and the costs.
{
"role": "user",
"content": [
{ "type": "text", "text": "What do these say, and does the clip match?" },
{ "type": "file", "file": { "file_id": "3f2a9c1e-7b4d-4e0a-9c1f-2d8e6a5b4c3d" } },
{ "type": "file", "file": { "filename": "notes.pdf", "file_data": "data:application/pdf;base64,JVBERi0xLjcK..." } },
{ "type": "input_audio", "input_audio": { "data": "UklGRiQAAABXQVZF...", "format": "wav" } },
{ "type": "video_url", "video_url": { "url": "data:video/mp4;base64,AAAAIGZ0eXBpc29t..." } }
]
}Cost. A file is input tokens at the model's input rate. The hold reserves 1,600 tokens a picture, 1,600 a PDF page, 32 a second of sound, 100 a second of video, and the text's own count where a file goes as words; it settles to the gateway's real figure. The three newest user messages that name uploads carry them in full; older ones get a one-line note in the file's place, so a PDF at the top of a long conversation is not paid for on every call.
Refused before anything runs. Nothing is reserved for these. Errors.
| Status | Message | Why |
|---|---|---|
| 400 | The model 'x' does not accept image input. | An inline picture on a model whose inputs lack image. |
| 400 | The model 'x' does not accept file input. Upload the file with POST /v1/files and send its file_id instead; askr then sends the model its text where it can. | An inline file, input_audio or video_url part on a model that takes no such input. By file_id, a PDF or sound goes as words instead. |
| 400 | A file part needs file.file_id (an upload), or file.file_data as a PDF data URL or an https URL. | The part is malformed. input_audio and video_url parts say the same in their own words. |
| 404 | No file with id 'x'. | The file_id names nothing on your account: deleted, swept after 7 days unused, or someone else's. |
| 413 | At most 6 pdf files per request. | More than the per-request count for that kind, across every message sent. Limits. |
| 413 | This request, with its files, is about 180,000 tokens; the model 'x' takes 128,000. | Over the model's window. code is context_length_exceeded. Pick a bigger window, or send less. |
| 400 | That model cannot watch video. Pick one that does, or remove the video. | A workspace code in code: W609 here; also W203, W605, W606, W607, W608 and W611. Error codes. |
Limits
- Requests
- 120 per minute per IP address
- Body size
- 12 MB on this endpoint, for inline pictures and files; uploads go to
/v1/filesat up to 100 MB - Output
max_tokensup to 8,192; default 4,096- Deadline
- 120 seconds upstream. Past it you get
504and the reservation is charged, because the model did run. Errors
Search the web
Send "web_search": true and the model can search the web and read pages during the call, exactly as the globe in the workspace does. The model is given two tools of ours, web_search and fetch_page, decides for itself whether to use them, and answers with numbered [n] marks. Per call: at most 3 searches, 2 pages read and 4 model calls, the last of them made without tools. It works on a model whose entry on GET /v1/models says supportsTools: true, streamed or not. Your own tools are still dropped.
curl https://heyaskr.ai/v1/chat/completions \
-H "Authorization: Bearer $ASKR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-5.6-sol",
"web_search": true,
"messages": [{"role": "user", "content": "When did askr add web search?"}]
}'{
"id": "chatcmpl-9c2b...",
"object": "chat.completion",
"created": 1789516800,
"model": "gpt-5.6-sol",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Search the web went live on 23 September [1].",
"annotations": [
{
"type": "url_citation",
"url_citation": { "url": "https://heyaskr.ai/changelog", "title": "Changelog", "start_index": 41, "end_index": 44 }
}
]
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 3412,
"completion_tokens": 58,
"total_tokens": 3470,
"searches": 1,
"fetches": 0,
"search_credits": 5
},
"askr": { "credits_charged": 5.31 }
}| Field | Meaning |
|---|---|
| message.annotations | One per source the call read, in the order the answer numbers them: [1] is annotations[0]. The shape is OpenAI's url_citation, so a client that reads OpenAI's search results reads ours. |
| url_citation.url, url_citation.title | The page, and its title. |
| url_citation.start_index, url_citation.end_index | The span of the first [n] mark in content, in UTF-16 code units as OpenAI counts them. A source the model listed but never cited inline has an empty span at the end of the text. |
| usage.searches, usage.fetches | How many searches ran and how many pages were read. Zero and zero when the model chose not to search. |
| usage.search_credits | The searches at 5 credits each. |
| askr.credits_charged | The whole charge: the model's tokens across every call plus search_credits. |
Streamed, the chunks are the ones on Streaming, with three differences: content from every model call arrives in order with a blank line between calls; a : askr comment line appears while a search or a page read runs, so an idle proxy does not close the stream; and the last data frame carries delta.annotations, finish_reason, usage and askr together.
data: {"id":"chatcmpl-9c2b...","object":"chat.completion.chunk","created":1789516800,"model":"gpt-5.6-sol","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
: askr
data: {"id":"chatcmpl-9c2b...","object":"chat.completion.chunk","created":1789516800,"model":"gpt-5.6-sol","choices":[{"index":0,"delta":{"content":"Search the web went live"},"finish_reason":null}]}
...
data: {"id":"chatcmpl-9c2b...","object":"chat.completion.chunk","created":1789516800,"model":"gpt-5.6-sol","choices":[{"index":0,"delta":{"content":"","annotations":[{"type":"url_citation","url_citation":{"url":"https://heyaskr.ai/changelog","title":"Changelog","start_index":41,"end_index":44}}]},"finish_reason":"stop"}],"usage":{"prompt_tokens":3412,"completion_tokens":58,"total_tokens":3470,"searches":1,"fetches":0,"search_credits":5},"askr":{"credits_charged":5.31}}
data: [DONE]Cost. Each search is 5 credits. The text the model reads is billed as input tokens at the model's rate, like the rest of the conversation, and every extra model call sends the conversation again. The hold is raised by 15 credits plus the most text the searches, the pages and the extra calls could add at the model's input rate; a 402 names that figure. It settles once, to the searches that ran plus the real token usage across every call. A call cut short is settled on what streamed, at most the hold. A call that ends with no content releases the hold and charges nothing, searches included. Each account may start 6 searching calls a minute, workspace and API together.
Refused before anything runs. Nothing is reserved for these. All three carry "param": "web_search" and the workspace's code in code, so one code means one thing on both surfaces. Errors.
| Status | Type | Message | Why |
|---|---|---|---|
| 400 | invalid_request_error | web_search is not available on this server. (code W206) | This server has no search. Drop web_search. |
| 400 | invalid_request_error | This model cannot use web_search; see supportsTools on GET /v1/models. (code W207) | The model takes no tools, or answers with pictures. Pick one with supportsTools: true. |
| 429 | rate_limit_error | Search the web is limited to a few turns a minute. Wait a moment and send again. (code W109) | More than 6 searching calls in a minute on the account. Wait and retry. |