Docs/API

Chat completions

Every field we read, every limit we apply, and what comes back.

Checked against the code on

Endpoint
POST https://heyaskr.ai/v1/chat/completions

Create a model response for a conversation. The request is the OpenAI shape. This page is precise about which fields we read, because the rest are dropped without a warning.

Request fields we read

FieldTypeWhat it does
modelstringA chat model id from the catalog. Exact match; there are no suffixes or aliases.
messagesarrayThe conversation. Each item has role (system, user or assistant) and content: a string, or an array of parts for pictures and files. Pictures and files.
max_tokensintegerOutput ceiling, thinking included. Default 4,096, maximum 8,192. Larger values are clamped, not rejected.
streambooleantrue for server-sent events. Streaming.
web_searchbooleanaskr's own field. true lets the model search the web and read pages during the call; the sources come back as annotations. Search the web.

Everything else is dropped silently: temperature, top_p, tools, tool_choice, response_format, stop, n, seed, plugins. A request that sends them succeeds; they just have no effect. Tool calling and structured output are on the roadmap. Web search is not a tool you define: it is the web_search field above.

Messages

  • Roles other than system, user and assistant return 400 Unsupported message role.
  • content is a string, or an array of parts: text, image_url, file, input_audio and video_url, each inline or by file_id. Any other part type returns 400 Unsupported content part. A message that is only parts is sent as it is. Pictures and files.
  • Only the last 40 messages are sent to the model. Trim on your side if you want control over what is kept.
  • A conversation may fill the model's own context window, less room for the reply: about 1 million tokens on the largest models, such as Kimi K3, at most 3.2 million characters and the last 400 messages. A model with a small window keeps 120,000 characters and 40 messages. Over the limit returns 400 The conversation is too long.
  • No usable messages returns 400 No usable messages were provided.

Example

curl https://heyaskr.ai/v1/chat/completions \
  -H "Authorization: Bearer $ASKR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5",
    "max_tokens": 300,
    "messages": [
      {"role": "system", "content": "Answer in one sentence."},
      {"role": "user", "content": "Why is the sky blue?"}
    ]
  }'

Response

200
{
  "id": "chatcmpl-9c2b...",
  "object": "chat.completion",
  "created": 1789516800,
  "model": "claude-sonnet-5",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Sunlight scatters off air molecules, and blue scatters most." },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 21,
    "completion_tokens": 14,
    "total_tokens": 35,
    "completion_tokens_details": { "reasoning_tokens": 0 }
  },
  "askr": { "credits_charged": 0.192 }
}
FieldMeaning
choices[0].message.contentThe reply. Always one choice.
choices[0].message.imagesPresent only when the model answered with a picture. Images from chat.
choices[0].message.annotationsPresent only on a web_search: true call: one url_citation per source. Search the web.
usage.completion_tokens_details.reasoning_tokensOn a thinking model, output tokens you never see. They are billed like any other output token.
askr.credits_chargedWhat the call cost, in credits, to four decimal places. Additive: clients that do not know it ignore it.

How the cost is worked out

Before the call, credits are reserved for the worst case: your estimated input tokens at the model's input rate, plus max_tokens at its output rate. When the model finishes, the reservation is settled to the real cost, using the gateway's own figure for the call. Nothing else is added: there is no platform fee in the first version. Holds and settlement.

Pictures and files in a message

A user message's content can be an array of parts, in OpenAI's shape. Five kinds are read. A part carries its bytes inline, or names an upload by file_id from POST /v1/files; an upload is the way to send anything big, or anything more than once.

PartInlineBy uploadInline limit
texttext
image_urlimage_url.url: a PNG, JPEG, WebP or GIF data URLimage_url.file_id8 MB decoded; 6 a message, 20 a request
filefile.file_data: a PDF as a data URL, or an https URL the provider fetches. file.filename is optional.file.file_id32 MB decoded
input_audioinput_audio.data as base64, with input_audio.format wav or mp3input_audio.file_id25 MB decoded
video_urlvideo_url.url: an MP4, MOV or WebM data URL, or an https URLvideo_url.file_id25 MB decoded

The request body on this endpoint is 12 MB, so an inline part over about 9 MB does not fit: upload it. An inline part goes to the model as it is, so it works only on a model whose inputs on GET /v1/models include that kind. A part by file_id is routed per model: the file itself where the model takes the kind, its words where it does not, so a PDF by file_id works on every model. Files has the rules and the costs.

{
  "role": "user",
  "content": [
    { "type": "text", "text": "What do these say, and does the clip match?" },
    { "type": "file", "file": { "file_id": "3f2a9c1e-7b4d-4e0a-9c1f-2d8e6a5b4c3d" } },
    { "type": "file", "file": { "filename": "notes.pdf", "file_data": "data:application/pdf;base64,JVBERi0xLjcK..." } },
    { "type": "input_audio", "input_audio": { "data": "UklGRiQAAABXQVZF...", "format": "wav" } },
    { "type": "video_url", "video_url": { "url": "data:video/mp4;base64,AAAAIGZ0eXBpc29t..." } }
  ]
}

Cost. A file is input tokens at the model's input rate. The hold reserves 1,600 tokens a picture, 1,600 a PDF page, 32 a second of sound, 100 a second of video, and the text's own count where a file goes as words; it settles to the gateway's real figure. The three newest user messages that name uploads carry them in full; older ones get a one-line note in the file's place, so a PDF at the top of a long conversation is not paid for on every call.

Refused before anything runs. Nothing is reserved for these. Errors.

StatusMessageWhy
400The model 'x' does not accept image input.An inline picture on a model whose inputs lack image.
400The model 'x' does not accept file input. Upload the file with POST /v1/files and send its file_id instead; askr then sends the model its text where it can.An inline file, input_audio or video_url part on a model that takes no such input. By file_id, a PDF or sound goes as words instead.
400A file part needs file.file_id (an upload), or file.file_data as a PDF data URL or an https URL.The part is malformed. input_audio and video_url parts say the same in their own words.
404No file with id 'x'.The file_id names nothing on your account: deleted, swept after 7 days unused, or someone else's.
413At most 6 pdf files per request.More than the per-request count for that kind, across every message sent. Limits.
413This request, with its files, is about 180,000 tokens; the model 'x' takes 128,000.Over the model's window. code is context_length_exceeded. Pick a bigger window, or send less.
400That model cannot watch video. Pick one that does, or remove the video.A workspace code in code: W609 here; also W203, W605, W606, W607, W608 and W611. Error codes.

Limits

Requests
120 per minute per IP address
Body size
12 MB on this endpoint, for inline pictures and files; uploads go to /v1/files at up to 100 MB
Output
max_tokens up to 8,192; default 4,096
Deadline
120 seconds upstream. Past it you get 504 and the reservation is charged, because the model did run. Errors

Search the web

Send "web_search": true and the model can search the web and read pages during the call, exactly as the globe in the workspace does. The model is given two tools of ours, web_search and fetch_page, decides for itself whether to use them, and answers with numbered [n] marks. Per call: at most 3 searches, 2 pages read and 4 model calls, the last of them made without tools. It works on a model whose entry on GET /v1/models says supportsTools: true, streamed or not. Your own tools are still dropped.

curl https://heyaskr.ai/v1/chat/completions \
  -H "Authorization: Bearer $ASKR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-5.6-sol",
    "web_search": true,
    "messages": [{"role": "user", "content": "When did askr add web search?"}]
  }'
200
{
  "id": "chatcmpl-9c2b...",
  "object": "chat.completion",
  "created": 1789516800,
  "model": "gpt-5.6-sol",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Search the web went live on 23 September [1].",
        "annotations": [
          {
            "type": "url_citation",
            "url_citation": { "url": "https://heyaskr.ai/changelog", "title": "Changelog", "start_index": 41, "end_index": 44 }
          }
        ]
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 3412,
    "completion_tokens": 58,
    "total_tokens": 3470,
    "searches": 1,
    "fetches": 0,
    "search_credits": 5
  },
  "askr": { "credits_charged": 5.31 }
}
FieldMeaning
message.annotationsOne per source the call read, in the order the answer numbers them: [1] is annotations[0]. The shape is OpenAI's url_citation, so a client that reads OpenAI's search results reads ours.
url_citation.url, url_citation.titleThe page, and its title.
url_citation.start_index, url_citation.end_indexThe span of the first [n] mark in content, in UTF-16 code units as OpenAI counts them. A source the model listed but never cited inline has an empty span at the end of the text.
usage.searches, usage.fetchesHow many searches ran and how many pages were read. Zero and zero when the model chose not to search.
usage.search_creditsThe searches at 5 credits each.
askr.credits_chargedThe whole charge: the model's tokens across every call plus search_credits.

Streamed, the chunks are the ones on Streaming, with three differences: content from every model call arrives in order with a blank line between calls; a : askr comment line appears while a search or a page read runs, so an idle proxy does not close the stream; and the last data frame carries delta.annotations, finish_reason, usage and askr together.

SSE
data: {"id":"chatcmpl-9c2b...","object":"chat.completion.chunk","created":1789516800,"model":"gpt-5.6-sol","choices":[{"index":0,"delta":{"role":"assistant","content":""},"finish_reason":null}]}
: askr
data: {"id":"chatcmpl-9c2b...","object":"chat.completion.chunk","created":1789516800,"model":"gpt-5.6-sol","choices":[{"index":0,"delta":{"content":"Search the web went live"},"finish_reason":null}]}
...
data: {"id":"chatcmpl-9c2b...","object":"chat.completion.chunk","created":1789516800,"model":"gpt-5.6-sol","choices":[{"index":0,"delta":{"content":"","annotations":[{"type":"url_citation","url_citation":{"url":"https://heyaskr.ai/changelog","title":"Changelog","start_index":41,"end_index":44}}]},"finish_reason":"stop"}],"usage":{"prompt_tokens":3412,"completion_tokens":58,"total_tokens":3470,"searches":1,"fetches":0,"search_credits":5},"askr":{"credits_charged":5.31}}
data: [DONE]

Cost. Each search is 5 credits. The text the model reads is billed as input tokens at the model's rate, like the rest of the conversation, and every extra model call sends the conversation again. The hold is raised by 15 credits plus the most text the searches, the pages and the extra calls could add at the model's input rate; a 402 names that figure. It settles once, to the searches that ran plus the real token usage across every call. A call cut short is settled on what streamed, at most the hold. A call that ends with no content releases the hold and charges nothing, searches included. Each account may start 6 searching calls a minute, workspace and API together.

Refused before anything runs. Nothing is reserved for these. All three carry "param": "web_search" and the workspace's code in code, so one code means one thing on both surfaces. Errors.

StatusTypeMessageWhy
400invalid_request_errorweb_search is not available on this server. (code W206)This server has no search. Drop web_search.
400invalid_request_errorThis model cannot use web_search; see supportsTools on GET /v1/models. (code W207)The model takes no tools, or answers with pictures. Pick one with supportsTools: true.
429rate_limit_errorSearch the web is limited to a few turns a minute. Wait a moment and send again. (code W109)More than 6 searching calls in a minute on the account. Wait and retry.