We tested every model on askr against its label.
Here is what we found.
278 confirmed as labelled, and no model answered like an older model than its label. The other 105 could not be confirmed either way: too old for the test, answering by web search or built as tools, not answering, or not settled yet. Every one is listed in section 4.
On 26 September a user published a test saying three models on askr were older models than their labels: Grok 4.6, Grok 4.7 and GLM 5.3. We ran their test, then sharper ones, on all 383 chat models we offer. The three models are the models on their labels. What made them look older was partly askr's doing and partly a supplier route's, and this page shows both.
Bug reports are welcome, in any form
A screenshot, a message in the Telegram group, a line through the feedback bar, a full write-up like this one. We thank everyone who tells us something is wrong, and we act on it. That is why the site is covered in prompts to give feedback, why you can book a call with the founders, and why support answers as fast as we can. We are building a consumer app that puts the user experience first, and that only works if you tell us where it breaks.
1. The bug and the test that found it
askr gives you a few hundred AI models under their makers' names. If you pick Grok 4.7, you should get Grok 4.7. On 26 September a user published a report titled "Is ASKR selling the AI model it says it is? No." They had tested three of our models over five days: grok-4.6, grok-4.7 and glm-5.3.
Their test is simple and anyone can run it. A model only knows the world up to the day its training data ends, so they asked each model about well-known events from late 2025 and early 2026, then asked the same models from official sources. The official copies knew the events. Ours said they had not happened.
What our Grok 4.7 said, and what the official one said
| Question | Official Grok 4.7 | Grok 4.7 on askr |
|---|---|---|
| When did Google release Gemini 3? | 18 November 2025 | "No confirmed release date" |
| When did OpenAI release GPT-5? | 7 August 2025 | "OpenAI has not announced GPT-5" |
| What did xAI release in November 2025? | Grok 4.1 | "I don't have confirmed information" |
| Name something from January 2026 | CES 2026 in Las Vegas | Could not name anything |
From the user's report. The official copy ran inside another app, with that app's own instructions. Ours ran with none, and without being told the date.
The report found a real bug. It was not the one it looked like: the model behind the label was the right one, answering as if its training had only just ended. Section 2 shows why. We are grateful to the user. This is the test we should have been running ourselves, and from now on we do.
2. What we found
We read our own code first. askr passes the model name you pick to our supplier unchanged, with your messages. In solo chat and on the API we add no instructions (only Search the web adds a line with its rules), we never fall back to a different model, and an unknown model name is refused, not replaced. We also add no date. That turned out to matter.
Then we asked the same three models the same kind of questions, one at a time, twice: once as the user did, and once with a single line in front saying what today's date is.
| Model on askr | Without the date | With the date | Verdict |
|---|---|---|---|
| Grok 4.7x-ai/grok-4.7 | "I don't have a confirmed winner" of the 2025 Nobel Peace Prize. The 2026 Winter Olympics "have not yet been held". Yet it named GLM-5's release on 11 February 2026. | María Corina Machado, 10 October 2025. Gemini 3 on 18 November 2025. GLM-5 on 11 February 2026. Asked to answer from its own knowledge: Claude Opus 4.6 on 5 February 2026, and Norway topping the 2026 medal table. | As labelled |
| Grok 4.6grok-4.6 | "Francis remains pope." "GPT-5 has not been released." Yet it described the 19 October 2025 Louvre theft in detail. | Pope Leo XIV. María Corina Machado, announced 10 October 2025. It still denied GPT-5 and misdated some later events. | As labelled |
| GLM 5.3glm-5.3 | "My knowledge has a cutoff in early 2025." Could not name the pope elected in May 2025, yet named the 2025 Nobel Peace Prize winner. | Pope Leo XIV on 8 May 2025. GPT-5 on 7 August 2025. María Corina Machado. Asked to answer from its own knowledge: Gemini 3 on 18 November 2025. | As labelled |
| GLM 5.3 by Engyengy/glm-5.3 | No 2025 Nobel winner, no Louvre theft, no Gemini 3. | María Corina Machado. Gemini 3 around 18 November 2025. Asked to answer from its own knowledge: Grok 4.1 on 17 November 2025. | As labelled |
A model trained on data that ends in 2024 cannot name the 2025 Nobel Peace Prize winner or a model released in February 2026, whatever you tell it. Grok 4, from July 2025, could not either. These models can. And the date line cannot invent knowledge: Claude Haiku 4.5, taken straight from Anthropic, whose knowledge ends in early 2025, got the same line and still could not name a single one of these events.
What made them look older:
- No date. A model does not know what day it is. Without a date it tends to assume it is still early in its training, and it treats anything later as not yet happened. This is our part: askr sent the models no date. Official apps usually do.
- Hidden instructions on one route. On the route our supplier uses for Grok 4.7, about 1,240 tokens of instructions sit in front of every message before it reaches the model. We measured it with askr's code out of the path: a bare "hi" sent straight to our supplier counts 1,243 prompt tokens. We did not put them there, and they make the model guarded: in one run it gave Gemini 3's release date and details of the October 2025 Louvre theft, then said no pope had been elected in May 2025.
- Grok denies rivals' launches. Grok models often say GPT-5 or Gemini 3.1 do not exist while naming later events correctly. The report's questions were mostly about rivals' launches.
- Models are poor judges of their own name. Claude Haiku 4.5, taken straight from Anthropic, calls itself Claude 3.5 Sonnet. A Grok calling itself something else proves nothing.
- No script on Grok 4.6. The report says it caught Grok 4.6's hidden instructions, including "a knowledge cutoff of 2024-10". When we measured our supplier's route for Grok 4.6, a bare "hi" counted 19 prompt tokens: nothing of that size sits in front of it. Asked to repeat instructions it does not have, a model will often write out the kind it was trained with.
- GLM's family blind spot is Z.ai's own. Our GLM 5.3 does not know GLM-5 or GLM-4.7, but neither does GLM 5.3 FlashX served by Z.ai itself, asked the same questions with the date. And GLM-4.6 came out on 30 September 2025, so a model that knows about 18 November 2025 is not GLM-4.6.
- Tools were our gap. Our API does not yet pass tool definitions to the model; the API docs say so, and tool calling is on the roadmap. A model that is never shown the weather tool cannot call it, so it types the call out or says it will check. That one is on us.
- Cheap does not mean fake. We pay our supplier its price for the named model on every request. The holder discount and the free credits come out of askr's pocket, not from swapping in cheaper models.
- Different answers on different days. Our supplier can send the same model to a different host from one request to the next. The receipts in section 3 will show which host answered.
Two more things in the report were askr's, not a model's. "The conversation is too long" is askr's own error message, word for word: until 24 September we capped a conversation at 120,000 characters, and since release 1.384 each model gets its own full context window, which is why a longer document went through later. "The model returned no content" is our refund path, shown when a model spends its whole budget thinking; in our own first pass 16% of calls ended that way at a small budget, and a larger one fixed nearly all of them.
3. What we are doing about it
Every point in the report has a line here. Each shows its real status, and nothing is marked done until it is live for every user.
Why we are giving so much back right now
This is also why we are giving away so many free credits right now, why anyone who hits a problem gets refunded, even a small one, and why someone answers feedback around the clock. We would rather hear about a problem the day it happens and fix it than read about it in a thread a week later.
4. The full audit of every model
Our catalog lists 383 chat models (the picker hides the ones that are not answering), so we tested all of them on 26 September, askr's code out of the path, no search, no tools. Four rounds:
- The user's quiz, word for word, on every model.
- Three open questions per model about events a few months before its maker's stated cutoff.
- For the 25 models that still looked old, the same questions again, with and without today's date.
- For every model, one question listing 23 well-known events from March 2023 to June 2026, with today's date given.
The newest event a model gets right is where its knowledge ends. We compare that with where its maker says it ends. To know what a genuine model looks like, we ran the same tests on five Claude models taken straight from Anthropic's own API: their knowledge ends 3 to 8 months before Anthropic's stated cutoff. So a model is "as labelled" when the gap is 9 months or less. We only suspect a model when the gap is a year or more with nothing closer anywhere in its answers, and every suspected model then went to three independent reviewers whose job was to prove it genuine.
No model on askr was found to be a different model from the one on its label. 55 could not be settled either way; they are marked "cannot tell yet" below and stay under the nightly check.
- As labelled
- Knows what its maker says it knows, to within 9 months.
- Consistent
- An older model whose knowledge ends before this test's events begin. Nothing it said contradicts the label.
- Cannot tell yet
- Answered too little, or too guardedly, for a verdict either way.
- Not as labelled
- Knowledge ends a year or more before its maker's stated cutoff, confirmed by three reviewers.
- Not testable
- Searches the web to answer, or is a tool (a classifier, a code-apply model), or did not answer at all.
All 383 models
| Model | Where our supplier says it ran | Newest thing it knew | Maker's stated cutoff | Verdict |
|---|---|---|---|---|
| Aion 3.5aion-labs/aion-3.5 | AionLabs | Gemini 3 and Opus 4.5, Nov 2025 | not published | As labelled |
| Aion 3.5 Miniaion-labs/aion-3.5-mini | AionLabs | Mamdani win, Nov 2025; thin 2026 items | not published | As labelled |
| Aion-2.0aion-labs/aion-2.0 | AionLabs | Llama 4, Apr 2025 | not published | As labelled |
| Aion-3.0aion-labs/aion-3.0 | AionLabs | GPT-5, Aug 2025 | not published | As labelled |
| Aion-3.0-Miniaion-labs/aion-3.0-mini | AionLabs | Llama 4, Apr 2025 | not published | As labelled |
| Aion-RP 1.0 (8B)aion-labs/aion-rp-llama-3.1-8b | AionLabs | nothing reliable | Dec 2023 | Cannot tell yetIts answers to dated questions were largely made up, so we could not check its knowledge. |
| Claude Fable 5anthropic/claude-fable-5 | Anthropic, supplier's own key | GPT-5, Aug 2025 | Jan 2026 | As labelled |
| Claude Fable 5.1claude-fable-5.1 | Anthropic, supplier's own key | Gemini 3.1 Pro (Feb 2026) | Jun 2026 | As labelled |
| Claude Fable 5.1direct/claude-fable-5-1 | Anthropic (direct) | World Cup opener draw (Dec 2025) | Jun 2026 | As labelled |
| Claude Fable Latest~anthropic/claude-fable-latest | Anthropic, supplier's own key | Gemini 3.1 Pro, Feb 2026 | Jun 2026 | As labelled |
| Claude Haiku 4.5claude-haiku-4.5 | Amazon Bedrock, supplier's own key | GPT-4o (May 2024) | Feb 2025 | As labelled |
| Claude Haiku 4.5direct/claude-haiku-4-5 | Anthropic (direct) | 2024 US election result (Nov 2024) | Feb 2025 | As labelled |
| Claude Haiku Latest~anthropic/claude-haiku-latest | Amazon Bedrock, supplier's own key | DeepSeek-V3, Dec 2024 | Feb 2025 | As labelled |
| Claude Opus 4.1anthropic/claude-opus-4.1 | Amazon Bedrock, supplier's own key | Claude 3 models, Mar 2024 | Jan 2025 | Cannot tell yetIt recalled events only up to early 2024, about 10 months before the maker's stated cutoff. |
| Claude Opus 4.5anthropic/claude-opus-4.5 | Amazon Bedrock, supplier's own key | Pope Leo XIV, May 2025 | May 2025 | As labelled |
| Claude Opus 4.6anthropic/claude-opus-4.6 | Amazon Bedrock, supplier's own key | Pope Leo XIV, May 2025 | May 2025 | As labelled |
| Claude Opus 4.7anthropic/claude-opus-4.7 | Amazon Bedrock, supplier's own key | Nobel Peace Prize, Oct 2025 | Jan 2026 | As labelled |
| Claude Opus 4.8anthropic/claude-opus-4.8 | Amazon Bedrock, supplier's own key | GPT-5, Aug 2025 | Jan 2026 | As labelled |
| Claude Opus 5anthropic/claude-opus-5 | Anthropic, supplier's own key | Claude Opus 4.6, Feb 2026 | May 2026 | As labelled |
| Claude Opus 5direct/claude-opus-5 | Anthropic (direct) | Gemini 3 and Grok 4.1 (Nov 2025) | May 2026 | As labelled |
| Claude Opus 5.5claude-opus-5.5 | Amazon Bedrock, supplier's own key | Claude Opus 4.6 (Feb 2026) | Jun 2026 | As labelled |
| Claude Opus 5.5direct/claude-opus-5-5 | Anthropic (direct) | Claude Opus 4.6 (Feb 2026) | Jun 2026 | As labelled |
| Claude Opus Latest~anthropic/claude-opus-latest | Amazon Bedrock, supplier's own key | Winter Olympics result, Feb 2026 | Jun 2026 | As labelled |
| Claude Sonnet 4anthropic/claude-sonnet-4 | Amazon Bedrock, supplier's own key | o1-preview, Sep 2024 | Jan 2025 | As labelled |
| Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | Amazon Bedrock, supplier's own key | Claude 3 models, Mar 2024 | Jan 2025 | Cannot tell yetIts answers on recent events were inconsistent, recalling early 2024 in one test and early 2025 in another. |
| Claude Sonnet 4.6anthropic/claude-sonnet-4.6 | Amazon Bedrock, supplier's own key | Pope Leo XIV, May 2025 | Aug 2025 | As labelled |
| Claude Sonnet 5claude-sonnet-5 | Amazon Bedrock, supplier's own key | OpenAI Sora 2 (Sep 2025) | Jan 2026 | As labelled |
| Claude Sonnet 5direct/claude-sonnet-5 | Anthropic (direct) | Louvre jewel theft (Oct 2025) | Jan 2026 | As labelled |
| Claude Sonnet Latest~anthropic/claude-sonnet-latest | Amazon Bedrock, supplier's own key | Claude Sonnet 4.5, Sep 2025 | Jan 2026 | As labelled |
| Codestral 2508mistralai/codestral-2508 | Mistral | Llama 3.1 405B, Jul 2024 | not published | As labelled |
| Command Acohere/command-a | Cohere | GPT-4o (May 2024) | not published | As labelled |
| Command A+cohere/command-a-plus | Cohere | Claude 3.5 Sonnet, Jun 2024 | not published | Cannot tell yetServed by Cohere itself, but it showed nothing newer than mid-2024 and Cohere gives no cutoff to test against. |
| Command R (08-2024)cohere/command-r-08-2024 | Cohere | GPT-4 Turbo at OpenAI DevDay, Nov 2023 | not published | As labelled |
| Command R+ (08-2024)cohere/command-r-plus-08-2024 | Cohere | GPT-4, Mar 2023 | not published | Cannot tell yetServed by Cohere itself; it says its knowledge ends in Jan 2023 and showed nothing newer than GPT-4. |
| Command R7B (12-2024)cohere/command-r7b-12-2024 | Cohere | GPT-4o, May 2024 | not published | As labelled |
| Cydonia 24B V4.1thedrummer/cydonia-24b-v4.1 | Parasail | Llama 3.1 405B, Jul 2024 | not published | As labelled |
| DeepSeek Flash Latest~deepseek/deepseek-flash-latest | Fireworks | Mamdani NYC mayor win, Nov 2025 | not published | As labelled |
| DeepSeek Pro Latest~deepseek/deepseek-pro-latest | Baidu | Gemini 3, late 2025 | not published | As labelled |
| DeepSeek V3deepseek/deepseek-chat | StreamLake | GPT-4o, May 2024 | not published | As labelled |
| DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | SiliconFlow | GPT-4o, May 2024 | not published | As labelled |
| DeepSeek V3.1deepseek/deepseek-chat-v3.1 | DeepInfra | Claude 3.5 Sonnet, Jun 2024 | not published | As labelled |
| DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | Novita | OpenAI o1-preview, Sep 2024 | not published | As labelled |
| DeepSeek V3.2deepseek/deepseek-v3.2 | Phala | OpenAI o1-preview, Sep 2024 | not published | As labelled |
| DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | AtlasCloud | OpenAI o1-preview, Sep 2024 | not published | As labelled |
| DeepSeek V4 Flash (Private via TEE)private/deepseek-v4-flash | not reported | no answer | not published | Not answeringBoth attempts returned unreadable garbage instead of a reply, so the model could not be tested. |
| DeepSeek V4 Flash 0423deepseek/deepseek-v4-flash | OpenInference | Meta Llama 4, Apr 2025 | not published | As labelled |
| DeepSeek V4 Flash 0731deepseek/deepseek-v4-flash-0731 | OpenInference | Claude Sonnet 4, May 2025 | not published | As labelled |
| DeepSeek V4 Flash Latest~deepseek/deepseek-v4-flash-latest | Relace | DeepSeek V3, Dec 2024 | not published | As labelled |
| DeepSeek V4 Flash Vision Expdeepseek/deepseek-v4-flash-vision-exp | SiliconFlow | Mamdani NYC mayor win, Nov 2025 | not published | As labelled |
| DeepSeek V4 Pro 0423deepseek/deepseek-v4-pro | StreamLake | Meta Llama 4, Apr 2025 | not published | As labelled |
| DeepSeek V4 Pro 0813deepseek/deepseek-v4-pro-0813 | Novita | Claude 3.7 Sonnet, Feb 2025 | not published | Cannot tell yetIts answers stopped in early 2025 or earlier, about 10 months short of what this model should know; not proven either way. |
| DeepSeek V4.1 Flashdeepseek/deepseek-v4.1-flash | not reported | GPT-5.2, Dec 2025 | not published | As labelled |
| DeepSeek V4.1 Flash (Private via TEE)private/deepseek-v4-1-flash | not reported | NYC mayoral result, Nov 2025 | not published | As labelled |
| deepseek-v4-flash-0731 by Engyengy/deepseek-v4-flash-0731 | not reported | Pope Leo XIV elected, May 2025 | not published | As labelled |
| deepseek-v4.1-flash by Engyengy/deepseek-v4.1-flash | not reported | Gemini 3 (Nov 2025) | not published | As labelled |
| Devstral 2 2512mistralai/devstral-2512 | Mistral | OpenAI o1-preview, Sep 2024 | not published | As labelled |
| Dots3-Note Previewdots-studio/dots-3-note-preview | not reported | no answer | not published | Not answeringWe could not run this model: no server was available to answer it. |
| Ember-1fireworks/ember-1 | not reported | Norway's Olympic medal lead, Feb 2026 | not published | As labelled |
| ERNIE 4.5 VL 424B A47Bbaidu/ernie-4.5-vl-424b-a47b | Novita | OpenAI o1-preview (Sep 2024) | not published | As labelled |
| Fugu Maxsakana/fugu-max | Sakana AI | World Cup opener draw, Dec 2025 | not published | As labelled |
| Fugu Ultrasakana/fugu-ultra | Sakana AI | 2025 Nobel Peace Prize, Oct 2025 | not published | As labelled |
| Fugu Ultra v2sakana/fugu-ultra-v2 | Sakana AI | GLM-5 release, Feb 2026 | Aug 2026 | As labelled |
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | not reported | Jul 2024: Llama 3.1 405B | Jan 2025 | As labelled |
| Gemini 2.5 Flash Litegoogle/gemini-2.5-flash-lite | not reported | May 2024: GPT-4o | Jan 2025 | As labelled |
| Gemini 2.5 Progoogle/gemini-2.5-pro | not reported | Nov 2024: Trump won US election | Jan 2025 | As labelled |
| Gemini 2.5 Pro Preview 06-05google/gemini-2.5-pro-preview | Google, supplier's own key | Nov 2024: Trump won US election | Jan 2025 | As labelled |
| Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | not reported | Jan 2025: DeepSeek-R1 | Jan 2025 | As labelled |
| Gemini 3.1 Flash Litegoogle/gemini-3.1-flash-lite | not reported | Feb 2025: Claude 3.7 Sonnet | Jan 2025 | As labelled |
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | Google AI Studio | Feb 2025: Claude 3.7 Sonnet | Jan 2025 | As labelled |
| Gemini 3.1 Pro Previewgoogle/gemini-3.1-pro-preview | not reported | Feb 2025: Claude 3.7 Sonnet | Jan 2025 | As labelled |
| Gemini 3.1 Pro Preview Custom Toolsgoogle/gemini-3.1-pro-preview-customtools | Google AI Studio | Feb 2025: Claude 3.7 Sonnet | Jan 2025 | As labelled |
| Gemini 3.5 Flashgoogle/gemini-3.5-flash | not reported | Feb 2025: Claude 3.7 Sonnet | Jan 2025 | As labelled |
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | not reported | Feb 2026: Norway top at Winter Olympics | Mar 2026 | As labelled |
| Gemini 3.6 Flashgoogle/gemini-3.6-flash | not reported | Feb 2026: Norway top at Winter Olympics | Mar 2026 | As labelled |
| Gemini 3.7 Flashgemini-3.7-flash | not reported | Feb 2025: Claude 3.7 Sonnet | Mar 2026 | Cannot tell yetKnew events to Feb 2025; its maker says some topics stop at Jan 2025, so this may be normal for it. |
| Gemini 3.8 Flashgoogle/gemini-3.8-flash | not reported | Feb 2026: Norway top at Winter Olympics | Mar 2026 | As labelled |
| Gemini Flash Latest~google/gemini-flash-latest | Claude 3.7 Sonnet, Feb 2025 | Mar 2026 | Cannot tell yetServed by Google itself; it knew little after early 2025, which Google's card allows for some topics. | |
| Gemini Pro Latest~google/gemini-pro-latest | Google, supplier's own key | Feb 2025: Claude 3.7 Sonnet | Jan 2025 | As labelled |
| Gemma 2 27Bgoogle/gemma-2-27b-it | NextBit | Nov 2023: OpenAI DevDay, GPT-4 Turbo | not published | As labelled |
| Gemma 3 12Bgoogle/gemma-3-12b-it | DeepInfra | May 2024: GPT-4o | Aug 2024 | As labelled |
| Gemma 3 27Bgoogle/gemma-3-27b-it | Parasail | May 2024: GPT-4o | Aug 2024 | As labelled |
| Gemma 3 4Bgoogle/gemma-3-4b-it | DeepInfra | no answer in this test | Aug 2024 | ConsistentNo answer: the model was rate-limited during this test. |
| Gemma 4 26B A4Bgoogle/gemma-4-26b-a4b-it | Darkbloom | no answer in this test | Jan 2025 | As labelled |
| Gemma 4 26B A4B Uncensoredvenice/e2ee-gemma-4-26b-a4b-uncensored-p | not reported | Claude 3.7 Sonnet, Feb 2025 | Jan 2025 | As labelled |
| Gemma 4 31Bgoogle/gemma-4-31b-it | DeepInfra | Feb 2025: Claude 3.7 Sonnet | Jan 2025 | As labelled |
| Gemma 4 31B (Private via TEE)private/gemma4-31b | not reported | Claude 3.7 Sonnet, Feb 2025 | Jan 2025 | As labelled |
| Gemma 4 Uncensoredvenice/gemma-4-uncensored | not reported | DeepSeek-R1, Jan 2025 | Jan 2025 | As labelled |
| GLM 4.5z-ai/glm-4.5 | Z.AI | DeepSeek-V3, Dec 2024 | not published | ConsistentIt ran out of room before answering this test; an earlier check fitted its expected age. |
| GLM 4.5 Airz-ai/glm-4.5-air | Novita | o1-preview, Sep 2024 | not published | As labelled |
| GLM 4.5Vz-ai/glm-4.5v | Novita | DeepSeek-R1, Jan 2025 | not published | As labelled |
| GLM 4.6z-ai/glm-4.6 | Z.AI | DeepSeek-R1, Jan 2025 | not published | As labelled |
| GLM 4.6Vz-ai/glm-4.6v | Z.AI | o1-preview, Sep 2024 | not published | As labelled |
| GLM 4.7z-ai/glm-4.7 | DeepInfra | Gemini 2.0 Flash, Dec 2024 | not published | As labelled |
| GLM 4.7 Flashz-ai/glm-4.7-flash | Cloudflare | US 'Liberation Day' tariffs, Apr 2025 | not published | As labelled |
| GLM 4.7 Flash Hereticvenice/olafangensan-glm-4.7-flash-heretic | not reported | Pope Leo XIV elected, May 2025 | not published | As labelled |
| GLM 5z-ai/glm-5 | Amazon Bedrock, supplier's own key | Claude Opus 4, May 2025 | not published | As labelled |
| GLM 5 Turboz-ai/glm-5-turbo | Z.AI | Pope Leo XIV elected, May 2025 | not published | As labelled |
| GLM 5.1z-ai/glm-5.1 | Phala | Claude 4 models, May 2025 | not published | As labelled |
| GLM 5.2z-ai/glm-5.2 | not reported | GPT-5, Aug 2025 | not published | As labelled |
| GLM 5.2 (Fast)glm-5.2-fast | Fireworks, supplier's own key | GPT-5, Aug 2025 | not published | As labelled |
| GLM 5.3glm-5.3 | not reported | Claude Opus 4.5, Nov 2025 | not published | As labelled |
| GLM 5.3 Flashz-ai/glm-5.3-flash | not reported | Gemini 3, Nov 2025 | not published | As labelled |
| GLM 5.3 FlashXz-ai/glm-5.3-flashx | Z.AI | Gemini 3 and Opus 4.5, Nov 2025 | not published | As labelled |
| GLM 5.3 Primez-ai/glm-5.3-prime | Alibaba | Claude Opus 4.5, Nov 2025 | not published | As labelled |
| GLM 5V Turboz-ai/glm-5v-turbo | Z.AI | Claude 3.7 Sonnet, Feb 2025 | not published | Cannot tell yetRecalls events only into early 2025, well before its late-2025 base; served by its maker, so likely genuine but thin on recent news. |
| GLM Flash Latest~z-ai/glm-flash-latest | Together | Gemini 3, Nov 2025 | not published | As labelled |
| GLM Latest~z-ai/glm-latest | Sail Research | US shutdown ended, Nov 2025 | not published | As labelled |
| glm-5.2 by Engyengy/glm-5.2 | not reported | GPT-5, Aug 2025 | not published | As labelled |
| GLM-5.3 (Private via TEE)private/glm-5-3 | not reported | NYC mayoral result, Nov 2025 | not published | As labelled |
| glm-5.3 by Engyengy/glm-5.3 | not reported | Gemini 3 and Grok 4.1, Nov 2025 | not published | As labelled |
| GLM-5.3 Flash (Private via TEE)private/glm-5-3-flash | not reported | NYC mayoral result, Nov 2025 | not published | As labelled |
| glm-5.3-flash by Engyengy/glm-5.3-flash | not reported | GPT-5.1 and Gemini 3 (Nov 2025) | not published | As labelled |
| GPT Astra Latest~openai/gpt-astra-latest | Amazon Bedrock, supplier's own key | Gemini 3.1 Pro, Feb 2026 | Apr 2026 | As labelled |
| GPT Audioopenai/gpt-audio | not reported | no answer | Oct 2023 | Not testableAudio-only model: it rejects text-only questions, so the text quiz cannot test it. |
| GPT Audio Miniopenai/gpt-audio-mini | not reported | no answer | Oct 2023 | Not testableAudio-only model: it rejects text-only questions, so the text quiz cannot test it. |
| GPT Chat Latestopenai/gpt-chat-latest | OpenAI | Claude Opus 4.8, May 2026 | Aug 2025 | As labelled |
| GPT Luna Latest~openai/gpt-luna-latest | Amazon Bedrock, supplier's own key | Winter Olympics medal table, Feb 2026 | May 2026 | As labelled |
| GPT Mini Latest~openai/gpt-mini-latest | not reported | no answer | Aug 2025 | Not answeringThe request was refused upstream on both tries, so no quiz ran. |
| GPT Sol Latest~openai/gpt-sol-latest | Amazon Bedrock, supplier's own key | Winter Olympics medal table, Feb 2026 | Apr 2026 | As labelled |
| GPT Terra Latest~openai/gpt-terra-latest | Amazon Bedrock, supplier's own key | World Cup draw, Dec 2025 | Feb 2026 | As labelled |
| GPT-3.5 Turboopenai/gpt-3.5-turbo | OpenAI | Nothing tested (all after its cutoff) | Sep 2021 | ConsistentIts stated cutoff (Sep 2021) predates every event we test, so this check cannot tell; nothing contradicts the label. |
| GPT-3.5 Turbo (older v0613)openai/gpt-3.5-turbo-0613 | Azure | OpenAI DevDay, GPT-4 Turbo, Nov 2023 | Sep 2021 | Cannot tell yetAnswers like a newer model: it knows late-2023 events, two years past this model's Sep 2021 cutoff. |
| GPT-3.5 Turbo 16kopenai/gpt-3.5-turbo-16k | OpenAI | OpenAI DevDay, GPT-4 Turbo, Nov 2023 | Sep 2021 | Cannot tell yetAnswers like a newer model: it knows late-2023 events, two years past this model's Sep 2021 cutoff. |
| GPT-3.5 Turbo Instructopenai/gpt-3.5-turbo-instruct | OpenAI | Nothing tested (all after its cutoff) | Sep 2021 | ConsistentIts stated cutoff (Sep 2021) predates every event we test, so this check cannot tell; nothing contradicts the label. |
| GPT-4openai/gpt-4 | Azure | OpenAI DevDay, GPT-4 Turbo, Nov 2023 | Dec 2023 | As labelled |
| GPT-4 Turboopenai/gpt-4-turbo | OpenAI | GPT-4 launch, Mar 2023 | Dec 2023 | As labelled |
| GPT-4.1openai/gpt-4.1 | OpenAI | GPT-4o, May 2024 | Jun 2024 | As labelled |
| GPT-4.1 Miniopenai/gpt-4.1-mini | OpenAI | OpenAI DevDay, GPT-4 Turbo, Nov 2023 | Jun 2024 | As labelled |
| GPT-4.1 Nanoopenai/gpt-4.1-nano | OpenAI | OpenAI DevDay, Nov 2023 (date off) | Jun 2024 | Cannot tell yetThis tiny model gave several confident wrong answers, so the test could not confirm its knowledge either way. |
| GPT-4oopenai/gpt-4o | OpenAI | OpenAI DevDay, GPT-4 Turbo, Nov 2023 | Oct 2023 | As labelled |
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | OpenAI | Claude 3 family, Mar 2024 | Oct 2023 | As labelled |
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | Azure | Claude 3 family, Mar 2024 | Oct 2023 | As labelled |
| GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | OpenAI | OpenAI DevDay, GPT-4 Turbo, Nov 2023 | Oct 2023 | As labelled |
| GPT-4o-miniopenai/gpt-4o-mini | OpenAI | OpenAI DevDay, GPT-4 Turbo, Nov 2023 | Oct 2023 | As labelled |
| GPT-4o-mini (2024-07-18)openai/gpt-4o-mini-2024-07-18 | OpenAI | OpenAI DevDay, GPT-4 Turbo, Nov 2023 | Oct 2023 | As labelled |
| GPT-5openai/gpt-5 | OpenAI | OpenAI o1-preview, Sep 2024 | Sep 2024 | As labelled |
| GPT-5 Miniopenai/gpt-5-mini | OpenAI | Claude 3 and GPT-4o, 2024 (dates off) | May 2024 | As labelled |
| GPT-5 Nanoopenai/gpt-5-nano | OpenAI | Claude 3 family, Mar 2024 | May 2024 | As labelled |
| GPT-5 Proopenai/gpt-5-pro | OpenAI | No answer | Sep 2024 | Not answeringIt used its whole output budget on reasoning and returned no answer, so we could not test it. |
| GPT-5.1openai/gpt-5.1 | OpenAI | OpenAI o1-preview, Sep 2024 | Sep 2024 | As labelled |
| GPT-5.1-Codexopenai/gpt-5.1-codex | Azure | Llama 3.1 405B, Jul 2024 | Sep 2024 | As labelled |
| GPT-5.1-Codex-Maxopenai/gpt-5.1-codex-max | Azure | GPT-4o, May 2024 | Sep 2024 | As labelled |
| GPT-5.1-Codex-Miniopenai/gpt-5.1-codex-mini | Azure | GPT-4o, 2024 (dates off) | Sep 2024 | As labelled |
| GPT-5.2openai/gpt-5.2 | OpenAI | DeepSeek-R1, Jan 2025 | Aug 2025 | As labelled |
| GPT-5.2 Chatopenai/gpt-5.2-chat | not reported | no answer | Aug 2025 | Not answeringThe model could not be reached: the maker has retired it, so the test could not run. |
| GPT-5.2 Proopenai/gpt-5.2-pro | OpenAI | DeepSeek-R1, Jan 2025 | Aug 2025 | As labelled |
| GPT-5.2-Codexopenai/gpt-5.2-codex | Azure | GPT-4o, May 2024 | Aug 2025 | Cannot tell yetOnly showed knowledge up to May 2024, over a year before its stated cutoff; its sibling models are also very cautious, so unproven. |
| GPT-5.3-Codexgpt-5.3-codex | OpenAI | GPT-4.1, Apr 2025 | Aug 2025 | As labelled |
| GPT-5.4openai/gpt-5.4 | Amazon Bedrock, supplier's own key | Grok 4 release, Jul 2025 | Aug 2025 | As labelled |
| GPT-5.4 Minigpt-5.4-mini | OpenAI | GPT-4.1, Apr 2025 | Aug 2025 | As labelled |
| GPT-5.4 Nanogpt-5.4-nano | OpenAI | DeepSeek-R1, Jan 2025 | Aug 2025 | As labelled |
| GPT-5.4 Proopenai/gpt-5.4-pro | OpenAI | Claude Opus 4, May 2025 | Aug 2025 | As labelled |
| GPT-5.5openai/gpt-5.5 | Amazon Bedrock, supplier's own key | Claude Opus 4.1, Aug 2025 | Dec 2025 | As labelled |
| GPT-5.5 Proopenai/gpt-5.5-pro | OpenAI | no answer returned | Dec 2025 | Not answeringIt spent its whole answer budget thinking and returned no text, so it could not be tested. |
| GPT-5.6 Lunaopenai/gpt-5.6-luna | Amazon Bedrock, supplier's own key | GPT-5.1, Nov 2025 | Feb 2026 | As labelled |
| GPT-5.6 Luna Proopenai/gpt-5.6-luna-pro | OpenAI | GPT-5.1, Nov 2025 | Feb 2026 | As labelled |
| GPT-5.6 Solopenai/gpt-5.6-sol | Amazon Bedrock, supplier's own key | World Cup draw, Dec 2025 | Feb 2026 | As labelled |
| GPT-5.6 Sol Proopenai/gpt-5.6-sol-pro | OpenAI | Mamdani wins NYC, Nov 2025 | Feb 2026 | As labelled |
| GPT-5.6 Terraopenai/gpt-5.6-terra | Amazon Bedrock, supplier's own key | Gemini 3, Nov 2025 | Feb 2026 | As labelled |
| GPT-5.6 Terra Proopenai/gpt-5.6-terra-pro | OpenAI | World Cup draw, Dec 2025 | Feb 2026 | As labelled |
| GPT-6 Astraopenai/gpt-6-astra | Amazon Bedrock, supplier's own key | Gemini 3.1 Pro, Feb 2026 | Apr 2026 | As labelled |
| GPT-6 Astra Progpt-6-astra-pro | OpenAI | Gemini 3.1 Pro, Feb 2026 | Apr 2026 | As labelled |
| GPT-6 Lunaopenai/gpt-6-luna | Amazon Bedrock, supplier's own key | Norway tops Olympics, Feb 2026 | May 2026 | As labelled |
| GPT-6 Luna Proopenai/gpt-6-luna-pro | OpenAI | Claude Opus 4.5, Nov 2025 | May 2026 | As labelled |
| GPT-6 Solgpt-6-sol | Amazon Bedrock, supplier's own key | Gemini 3 and NYC mayor result, Nov 2025 | Apr 2026 | As labelled |
| GPT-6 Sol Proopenai/gpt-6-sol-pro | OpenAI | Gemini 3.1 Pro, Feb 2026 | Apr 2026 | As labelled |
| GPT-OSS 120B (Private via TEE)private/gpt-oss-120b | not reported | GPT-4o, May 2024 | Jun 2024 | As labelled |
| gpt-oss-120bopenai/gpt-oss-120b | Amazon Bedrock, supplier's own key | GPT-4o, May 2024 | Jun 2024 | As labelled |
| gpt-oss-20bopenai/gpt-oss-20b | Amazon Bedrock, supplier's own key | GPT-4o, 2024 | Jun 2024 | As labelled |
| gpt-oss-safeguard-20bopenai/gpt-oss-safeguard-20b | Groq | GPT-4o, May 2024 | not published | Not testableA safety classifier, not a chat model, so a knowledge quiz does not apply; its answers match its base model. |
| Granite 4.0 Microibm-granite/granite-4.0-h-micro | Cloudflare | none reliable (dates all wrong) | not published | Cannot tell yetA very small model that gave wrong dates even for 2023 events, so its knowledge could not be pinned down. |
| Granite 4.2 8Bibm-granite/granite-4.2-8b | DeepInfra | May 2024: GPT-4o | not published | Cannot tell yetKnew events to about May 2024; its maker publishes no cutoff, and a sibling model looks the same. |
| Grok 4.20x-ai/grok-4.20 | xAI | DeepSeek-R1, Jan 2025 | not published | As labelled |
| Grok 4.20 Multi-Agentx-ai/grok-4.20-multi-agent | xAI | Pope Leo XIV elected, May 2025 | not published | As labelled |
| Grok 4.3x-ai/grok-4.3 | xAI | DeepSeek-R1, Jan 2025 | not published | Cannot tell yetKnows events only to Jan 2025, older than a reported Dec 2025 cutoff; xAI has not published one, so we can't confirm. |
| Grok 4.5x-ai/grok-4.5 | xAI | Claude Opus 4.6, Feb 2026 | not published | As labelled |
| Grok 4.6grok-4.6 | Amazon Bedrock, supplier's own key | Gemini 3 launch, Nov 2025 | not published | As labelled |
| Grok 4.7x-ai/grok-4.7 | xAI | Gemini 3.1 Pro, Feb 2026 | May 2026 | As labelled |
| Grok Build 0.1x-ai/grok-build-0.1 | xAI | Pope Leo XIV elected, May 2025 | not published | As labelled |
| Grok Latest~x-ai/grok-latest | xAI | Gemini 3.1 Pro, Feb 2026 | May 2026 | As labelled |
| Hermes 3 405B Instructnousresearch/hermes-3-llama-3.1-405b | DeepInfra | GPT-4 Turbo, Nov 2023 | Dec 2023 | As labelled |
| Hermes 3 70B Instructnousresearch/hermes-3-llama-3.1-70b | DeepInfra | Llama 2, Jul 2023 | Dec 2023 | As labelled |
| Hermes 4 405Bnousresearch/hermes-4-405b | Nebius | Llama 3.1, Jul 2024 | Dec 2023 | As labelled |
| Hunyuan A13B Instructtencent/hunyuan-a13b-instruct | SiliconFlow | Llama 3.1 405B, Jul 2024 | not published | As labelled |
| Hy-MT2-1.8Btencent/hy-mt2-1.8b | Tencent | none shown | not published | Not testableTranslation-only model, so a general-knowledge test does not apply. |
| Hy-MT2-30B-A3Btencent/hy-mt2-30b-a3b | Tencent | GPT-5 date, Aug 2025 (bare date only) | not published | Not testableTranslation-only model, so a general-knowledge test does not apply. |
| Hy-MT2-7Btencent/hy-mt2-7b | Tencent | Llama 2, Jul 2023 | not published | Not testableTranslation-only model, so a general-knowledge test does not apply. |
| Hy3tencent/hy3 | Phala | DeepSeek-V3, Dec 2024 | not published | Cannot tell yetKnew events up to Dec 2024; the maker gives no cutoff, so we cannot tell whether that is normal for it. |
| Hy3 previewtencent/hy3-preview | GMICloud | Llama 4, Apr 2025 | not published | As labelled |
| Hy4 previewtencent/hy4-preview | Novita | Gemini 3, Nov 2025 | not published | As labelled |
| Inklingthinkingmachines/inkling | DeepInfra | Norway tops Winter Olympics, Feb 2026 | not published | As labelled |
| Inkling Smallthinkingmachines/inkling-small | DeepInfra | NYC mayoral result, Nov 2025 | not published | As labelled |
| KAT-Coder-Pro V2.5kwaipilot/kat-coder-pro-v2.5 | not reported | not tested | not published | Not answeringThe model returned an error, so it could not be tested. |
| Kimi K2 0711moonshotai/kimi-k2 | Novita | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| Kimi K2 0905moonshotai/kimi-k2-0905 | Novita | Pope Leo XIV elected, May 2025 | not published | As labelled |
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | Novita | Llama 3.1, Jul 2024 | not published | Cannot tell yetShowed older knowledge than expected for this model; not conclusive, so we are rechecking it. |
| Kimi K2.5moonshotai/kimi-k2.5 | Amazon Bedrock, supplier's own key | Pope Leo XIV elected, May 2025 | not published | As labelled |
| Kimi K2.6moonshotai/kimi-k2.6 | Decart | Pope Leo XIV elected, May 2025 | not published | As labelled |
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | Moonshot AI | Pope Leo XIV elected, May 2025 | not published | As labelled |
| Kimi K3moonshotai/kimi-k3 | not reported | Kimi K2 Thinking, Nov 2025 | not published | As labelled |
| Kimi K3 (Fast)kimi-k3-fast | Fireworks, supplier's own key | Winter Olympics medal table, Feb 2026 | not published | As labelled |
| Kimi K3 (Private via TEE)private/kimi-k3 | not reported | NYC mayoral result, Nov 2025 | not published | As labelled |
| Kimi Latest~moonshotai/kimi-latest | Fireworks, supplier's own key | Kimi K2 Thinking, Nov 2025 | not published | As labelled |
| kimi-k3 by Engyengy/kimi-k3 | not reported | Kimi K2 Thinking (Nov 2025) | not published | As labelled |
| Laguna S 2.1poolside/laguna-s-2.1 | Poolside | Pope Leo XIV (Prevost), May 2025 | not published | As labelled |
| Laguna XS 2.1poolside/laguna-xs-2.1 | Poolside | Llama 4 and Pope Leo XIV, 2025 | not published | As labelled |
| LFM2.5-2.6Bliquid/lfm-2.5-2.6b | not reported | no answer | not published | Not answeringWe could not run this model: no server was available to answer it. |
| Ling 3.0 Flashinclusionai/ling-3.0-flash | DeepInfra | DeepSeek-R1, Jan 2025 | not published | As labelled |
| Ling 3.0 Flash Fininclusionai/ling-3.0-flash-fin | DeepInfra | US tariffs, Apr 2025 | not published | As labelled |
| Ling 3.0 Flash Santeinclusionai/ling-3.0-flash-sante | not reported | no answer | not published | Not answeringWe could not run this model: no server was available to answer it. |
| Ling 3.0 Flash VLinclusionai/ling-3.0-flash-vl | Novita | Pope Leo XIV elected, May 2025 | not published | Cannot tell yetKnowledge ends around mid-2025, a year before the maker's stated data date, probably because its text side is older. |
| Llama 3 8B Lunarissao10k/l3-lunaris-8b | Parasail | nothing on the list | Mar 2023 | ConsistentBuilt on a model whose knowledge ends in early 2023, so recent events are out of its range. |
| Llama 3.1 70B Instructmeta-llama/llama-3.1-70b-instruct | Amazon Bedrock, supplier's own key | GPT-4 and Llama 2, 2023 | Dec 2023 | As labelled |
| Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | DeepInfra | GPT-4, Mar 2023 | Dec 2023 | As labelled |
| Llama 3.1 Euryale 70B v2.2sao10k/l3.1-euryale-70b | DeepInfra | DevDay and GPT-4 Turbo, Nov 2023 | Dec 2023 | As labelled |
| Llama 3.2 1B Instructmeta-llama/llama-3.2-1b-instruct | Cloudflare | nothing reliable | Dec 2023 | Cannot tell yetA very small model; it guesses dates, so its answers about events cannot be trusted. |
| Llama 3.2 3B Instructmeta-llama/llama-3.2-3b-instruct | Parasail | OpenAI DevDay, Nov 2023 | Dec 2023 | As labelled |
| Llama 3.3 70B (Private via TEE)private/llama3-3-70b | not reported | Llama 2, Jul 2023 | Dec 2023 | As labelled |
| Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | Cloudflare | GPT-4 and Llama 2, 2023 | Dec 2023 | As labelled |
| Llama 3.3 Euryale 70Bsao10k/l3.3-euryale-70b | NextBit | DevDay and GPT-4 Turbo, Nov 2023 | Dec 2023 | As labelled |
| Llama 4 Maverickmeta-llama/llama-4-maverick | DeepInfra | OpenAI o1-preview, Sep 2024 | Aug 2024 | As labelled |
| Llama 4 Scoutmeta-llama/llama-4-scout | DeepInfra | Llama 3.1 405B, Jul 2024 | Aug 2024 | As labelled |
| Llama Guard 4 12Bmeta-llama/llama-guard-4-12b | DeepInfra | did not say | not published | Not testableA safety classifier that only replies safe or unsafe, so its knowledge cannot be tested. |
| LongCat 2.0meituan/longcat-2.0 | AtlasCloud | Claude Opus 4.5, Nov 2025 | not published | As labelled |
| Magnum v4 72Banthracite-org/magnum-v4-72b | Mancer 2 | GPT-4 Turbo, Nov 2023 | not published | Cannot tell yetIt gave the same made-up date for most questions, so we could not check its knowledge. |
| Mercury 2inception/mercury-2 | Inception | GPT-4o, May 2024 | not published | Cannot tell yetIts knowledge stops around mid-2024, early for a 2026 model; the maker publishes no cutoff, so we cannot confirm. |
| Mercury 2.5inception/mercury-2.5 | Inception | Claude 3.7 Sonnet, Feb 2025 | not published | Cannot tell yetIts knowledge stops around early 2025, early for a late-2026 model; the maker publishes no cutoff. |
| MiMo-V2.5xiaomi/mimo-v2.5 | Xiaomi | Louvre jewel theft, Oct 2025 | not published | As labelled |
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | GMICloud | Pope Leo XIV elected, May 2025 | not published | As labelled |
| MiMo-V2.6-Flashxiaomi/mimo-v2.6-flash | Xiaomi | Winter Olympics result, Feb 2026 | not published | As labelled |
| MiMo-V2.6-Proxiaomi/mimo-v2.6-pro | DeepInfra | Gemini 3, Nov 2025 | not published | As labelled |
| MiMo-V2.6-Pro-UltraSpeedxiaomi/mimo-v2.6-pro-ultraspeed | Xiaomi | GPT-5, Aug 2025 | not published | As labelled |
| MiniMax M1minimax/minimax-m1 | Minimax | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| MiniMax M2minimax/minimax-m2 | Minimax | Claude 3 family, Mar 2024 | not published | Cannot tell yetOn this test it said it knew nothing after early 2024, but it was very cautious; a follow-up is needed before drawing conclusions. |
| MiniMax M2-herminimax/minimax-m2-her | Minimax | Pope Leo XIV elected, May 2025 | not published | As labelled |
| MiniMax M2.1minimax/minimax-m2.1 | Minimax | Llama 3.1 405B, Jul 2024 | not published | Cannot tell yetShowed nothing after mid-2024, even on MiniMax's own servers; MiniMax publishes no cutoff to check it against. |
| MiniMax M2.5minimax/minimax-m2.5 | Friendli | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| MiniMax M2.7minimax/minimax-m2.7 | DeepInfra | Trump election win, Nov 2024 | not published | Cannot tell yetShowed nothing after late 2024 and denied Pope Leo XIV exists; MiniMax publishes no cutoff to check it against. |
| MiniMax M3minimax/minimax-m3 | StreamLake | Claude Sonnet 4.5, Sep 2025 | not published | As labelled |
| MiniMax-01minimax/minimax-01 | Minimax | GPT-4o, May 2024 | not published | As labelled |
| Ministral 3 14B 2512mistralai/ministral-14b-2512 | Mistral | GPT-4o, May 2024 | not published | As labelled |
| Ministral 3 3B 2512mistralai/ministral-3b-2512 | Mistral | GPT-4o, May 2024 | not published | As labelled |
| Ministral 3 8B 2512mistralai/ministral-8b-2512 | Mistral | GPT-4o, May 2024 | not published | As labelled |
| Mistral Largemistralai/mistral-large | Mistral | Gemini 1.5 Pro, Feb 2024 (earlier test) | not published | ConsistentThe date test was rate-limited and got no answer; an earlier quiz showed early-2024 knowledge. |
| Mistral Large 2407mistralai/mistral-large-2407 | Mistral | DeepSeek-V3, Dec 2024 | Oct 2023 | As labelled |
| Mistral Large 3 2512mistralai/mistral-large-2512 | Mistral | Gemini 1.5 Pro, Feb 2024 (earlier test) | not published | ConsistentThe date test was rate-limited and got no answer; an earlier quiz showed only early-2024 knowledge, so it needs a retest. |
| Mistral Medium 3mistralai/mistral-medium-3 | Mistral | OpenAI o1-preview, Sep 2024 | not published | As labelled |
| Mistral Medium 3.1mistralai/mistral-medium-3.1 | Mistral | OpenAI o1-preview, Sep 2024 | not published | As labelled |
| Mistral Medium 3.5mistralai/mistral-medium-3-5 | Mistral | Mistral Large 2, Jul 2024 | not published | Cannot tell yetServed by Mistral itself but showed nothing after mid-2024 for an April 2026 model; Mistral publishes no cutoff. |
| Mistral Nemomistralai/mistral-nemo | DeepInfra | Llama 2, Jul 2023 | not published | As labelled |
| Mistral Small 3mistralai/mistral-small-24b-instruct-2501 | DeepInfra | Claude 3 models, Mar 2024 | Oct 2023 | As labelled |
| Mistral Small 3.1 24Bmistralai/mistral-small-3.1-24b-instruct | Cloudflare | Claude 3 models, Mar 2024 | Oct 2023 | As labelled |
| Mistral Small 3.2 24Bmistralai/mistral-small-3.2-24b-instruct | DeepInfra | GPT-4o, May 2024 | Oct 2023 | As labelled |
| Mistral Small 4mistralai/mistral-small-2603 | Mistral | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| Mixtral 8x22B Instructmistralai/mixtral-8x22b-instruct | Mistral | nothing usable yet | not published | ConsistentThe date test was rate-limited and got no answer, so this model has not been checked yet. |
| Morph V3 Fastmorph/morph-v3-fast | Morph | did not say | not published | Not testableA code-editing model, not a chat model: it echoed the questions back, so a knowledge quiz does not apply. |
| Morph V3 Largemorph/morph-v3-large | Morph | did not say | not published | Not testableA code-editing model, not a chat model: it did not engage with the questions, so a knowledge quiz does not apply. |
| Muse Glimmer 30Bmeta/muse-glimmer-30b | Together | GPT-5, Aug 2025 | Jan 2026 | As labelled |
| Muse Spark 1.1meta/muse-spark-1.1 | Meta | Mamdani elected NYC mayor, Nov 2025 | not published | As labelled |
| Muse Spark 1.2meta/muse-spark-1.2 | Meta | Mamdani elected NYC mayor, Nov 2025 | not published | As labelled |
| Muse Spark 1.2 Contributormeta/muse-spark-1.2-contributor | not reported | no answer | not published | Not answeringBlocked: this tier trains on your data, and our data policy stops requests from reaching it. |
| Muse Spark 1.3meta/muse-spark-1.3 | Meta | Mamdani elected NYC mayor, Nov 2025 | not published | As labelled |
| Muse Spark 1.3 Contributormeta/muse-spark-1.3-contributor | not reported | no answer | not published | Not answeringBlocked: this tier trains on your data, and our data policy stops requests from reaching it. |
| MythoMax 13Bgryphe/mythomax-l2-13b | Parasail | nothing on this list (all 2023+) | Sep 2022 | ConsistentBuilt on a 2022-era base model; too old for this test to check. |
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | Nebius | o1-preview, Sep 2024 | Jun 2025 | As labelled |
| Nemotron 3 Nano Omninvidia/nemotron-3-nano-omni-30b-a3b-reasoning | not reported | no answer | not published | Not answeringNo paid route for this model id was available when we tested, so the quiz could not run. |
| Nemotron 3 Supernvidia/nemotron-3-super-120b-a12b | DekaLLM | GPT-4o, May 2024 | Jun 2025 | Cannot tell yetShowed about a year less knowledge than its maker states; not conclusive, so we are rechecking it. |
| Nemotron 3 Ultranvidia/nemotron-3-ultra-550b-a55b | BaseTen | DeepSeek-R1, Jan 2025 | Sep 2025 | As labelled |
| Nemotron 3.5 Content Safetynvidia/nemotron-3.5-content-safety | DeepInfra | did not say | not published | Not testableA safety classifier, not a chat model: it only labels messages safe or unsafe, so a knowledge quiz does not apply. |
| Nemotron 3.5 Lightningnvidia/nemotron-3.5-lightning | Io Net | o1-preview, Sep 2024 | Sep 2025 | Cannot tell yetShowed about a year less knowledge than its maker states; not conclusive, so we are rechecking it. |
| North Mini Codecohere/north-mini-code | not reported | no answer | not published | Not answeringNo paid version of this model is being served, so the request failed; it should be removed from the list. |
| Nova 2 Liteamazon/nova-2-lite-v1 | Amazon Bedrock, supplier's own key | DeepSeek V2, May 2024 | Oct 2025 | Cannot tell yetKnew nothing after mid-2024 and invented later events, though Amazon states an Oct 2025 cutoff; may be an older Nova. |
| Nova Lite 1.0amazon/nova-lite-v1 | Amazon Bedrock, supplier's own key | GPT-4o, May 2024 | Oct 2024 | As labelled |
| Nova Micro 1.0amazon/nova-micro-v1 | Amazon Bedrock, supplier's own key | Llama 2, Jul 2023 | Oct 2024 | Cannot tell yetIts answers were thin and partly wrong, so we could not confirm how recent its knowledge is. |
| Nova Premier 1.0amazon/nova-premier-v1 | not reported | no answer | Oct 2024 | Not answeringThe model has been retired by its maker and returned an error; it should be removed from the list. |
| Nova Pro 1.0amazon/nova-pro-v1 | Amazon Bedrock, supplier's own key | GPT-4 Turbo, Nov 2023 | Oct 2024 | Cannot tell yetIt recalled events only up to late 2023, about 11 months before Amazon's stated cutoff. |
| o1openai/o1 | OpenAI | GPT-4 Turbo, Nov 2023 | Oct 2023 | As labelled |
| o1-proopenai/o1-pro | OpenAI | GPT-4 Turbo, Nov 2023 | Oct 2023 | As labelled |
| o3openai/o3 | OpenAI | GPT-4o, May 2024 | Jun 2024 | As labelled |
| o3 Miniopenai/o3-mini | OpenAI | GPT-4 Turbo, Nov 2023 | Oct 2023 | As labelled |
| o3 Mini Highopenai/o3-mini-high | OpenAI | GPT-4 Turbo, Nov 2023 | Oct 2023 | As labelled |
| o3 Proopenai/o3-pro | OpenAI | GPT-4o, May 2024 | Jun 2024 | As labelled |
| o4 Miniopenai/o4-mini | OpenAI | GPT-4o, 2024 | Jun 2024 | As labelled |
| o4 Mini Highopenai/o4-mini-high | OpenAI | GPT-4o, 2024 | Jun 2024 | As labelled |
| Palmyra X5writer/palmyra-x5 | Amazon Bedrock, supplier's own key | Llama 3.1 405B, Jul 2024 | not published | As labelled |
| Paretounbiased/pareto | Unbiased | World Cup opener pairing, Dec 2025 | not published | As labelled |
| Perceptron Mk1perceptron/perceptron-mk1 | Perceptron | Llama 4 and Pope Leo XIV, 2025 | not published | As labelled |
| Perceptron Mk1.5perceptron/perceptron-mk1.5 | Perceptron | Mamdani NYC win, Nov 2025 | not published | As labelled |
| Phi 4microsoft/phi-4 | DeepInfra | Claude 3 family, 2024 | Jun 2024 | As labelled |
| Qwen Plus 0728qwen/qwen-plus-2025-07-28 | Alibaba | o1-preview, Sep 2024 | not published | As labelled |
| Qwen-Plusqwen/qwen-plus | Alibaba | o1-preview, Sep 2024 | not published | As labelled |
| Qwen2.5 72B Instructqwen/qwen-2.5-72b-instruct | DeepInfra | OpenAI DevDay, late 2023 | not published | As labelled |
| Qwen2.5 7B Instructqwen/qwen-2.5-7b-instruct | Phala | nothing shown | not published | Cannot tell yetRefused every question, even GPT-4, so this test says nothing about which model it is. |
| Qwen2.5 Coder 32B Instructqwen/qwen-2.5-coder-32b-instruct | Cloudflare | OpenAI DevDay, late 2023 | not published | As labelled |
| Qwen2.5 VL 72B Instructqwen/qwen2.5-vl-72b-instruct | Parasail | GPT-4o, May 2024 | not published | As labelled |
| Qwen3 14Bqwen/qwen3-14b | NextBit | DeepSeek-V3, late 2024 | not published | As labelled |
| Qwen3 235B A22Bqwen/qwen3-235b-a22b | Alibaba | DeepSeek-V3, late 2024 | not published | As labelled |
| Qwen3 235B A22B Instruct 2507qwen/qwen3-235b-a22b-2507 | DeepInfra | o1-preview, 2024 | not published | As labelled |
| Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | Novita | Llama 4, Apr 2025 | not published | As labelled |
| Qwen3 30B A3Bqwen/qwen3-30b-a3b | DeepInfra | Llama 2, Jul 2023 (later answers vague) | not published | Cannot tell yetAnswered with years only, often wrong, so this test cannot date it. |
| Qwen3 30B A3B Instruct 2507qwen/qwen3-30b-a3b-instruct-2507 | StreamLake | US election result, Nov 2024 | not published | As labelled |
| Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | Alibaba | Trump wins US election, Nov 2024 | not published | As labelled |
| Qwen3 32Bqwen/qwen3-32b | DeepInfra | o1-preview, Sep 2024 | not published | As labelled |
| Qwen3 8Bqwen/qwen3-8b | Alibaba | o1-preview, Sep 2024 | not published | As labelled |
| Qwen3 Coder 30B A3B Instructqwen/qwen3-coder-30b-a3b-instruct | SiliconFlow | no usable answer | not published | ConsistentThis model gave no usable answer to our knowledge test, so we could not check it. |
| Qwen3 Coder 480B A35Bqwen/qwen3-coder | DeepInfra | DeepSeek-R1, Jan 2025 | not published | As labelled |
| Qwen3 Coder Flashqwen/qwen3-coder-flash | Alibaba | Trump beats Harris, Nov 2024 | not published | As labelled |
| Qwen3 Coder Nextqwen/qwen3-coder-next | Alibaba | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| Qwen3 Coder Plusqwen/qwen3-coder-plus | Alibaba | Trump wins US election, Nov 2024 | not published | As labelled |
| Qwen3 Maxqwen/qwen3-max | Alibaba | Qwen3 release, Apr 2025 | not published | As labelled |
| Qwen3 Max Thinkingqwen/qwen3-max-thinking | Alibaba | Qwen3 release, Apr 2025 | not published | As labelled |
| Qwen3 Next 80B A3B Instructqwen/qwen3-next-80b-a3b-instruct | Alibaba | o1-preview, Sep 2024 | not published | As labelled |
| Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | Trump wins US election, Nov 2024 | not published | As labelled | |
| Qwen3 VL 235B A22B Instructqwen/qwen3-vl-235b-a22b-instruct | Parasail | o1-preview, Sep 2024 | not published | As labelled |
| Qwen3 VL 235B A22B Thinkingqwen/qwen3-vl-235b-a22b-thinking | Alibaba | Claude 3.5 Sonnet, Jun 2024 | not published | As labelled |
| Qwen3 VL 30B A3B Instructqwen/qwen3-vl-30b-a3b-instruct | DeepInfra | Trump wins US election, Nov 2024 | not published | As labelled |
| Qwen3 VL 30B A3B Thinkingqwen/qwen3-vl-30b-a3b-thinking | Alibaba | Rebels take Aleppo, Nov 2024 | not published | As labelled |
| Qwen3 VL 32B Instructqwen/qwen3-vl-32b-instruct | Alibaba | Trump wins US election, Nov 2024 | not published | As labelled |
| Qwen3 VL 8B Instructqwen/qwen3-vl-8b-instruct | Parasail | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| Qwen3 VL 8B Thinkingqwen/qwen3-vl-8b-thinking | Alibaba | DeepSeek-R1, Jan 2025 | not published | As labelled |
| Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | Phala | Llama 3, Apr 2024 | not published | Cannot tell yetIts answers stop in early 2024, well before its likely cutoff. The maker's own version answers the same way. |
| Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | Alibaba | o1-preview, Sep 2024 | not published | Cannot tell yetIt showed knowledge only to late 2024; its maker publishes no cutoff, so we cannot say whether that is expected. |
| Qwen3.5 Plus 2026-04-20qwen/qwen3.5-plus-20260420 | Alibaba | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | Novita | Qwen2.5 release, Sep 2024 | not published | Cannot tell yetIts answers stop around 2024, well before its likely cutoff. The maker's own version answers the same way. |
| Qwen3.5-27Bqwen/qwen3.5-27b | Novita | GPT-4o, May 2024 | not published | Cannot tell yetIts answers stop in mid-2024, well before its likely cutoff. The maker's own version answers the same way. |
| Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | Darkbloom | GPT-4o, May 2024 | not published | Cannot tell yetIts answers stop in mid-2024, well before its likely cutoff. This comes straight from the maker, so it is how the model answers. |
| Qwen3.5-9Bqwen/qwen3.5-9b | Darkbloom | no answer given | not published | Not answeringIt spent its whole answer budget thinking and gave no reply, so it could not be tested. |
| Qwen3.5-Flashqwen/qwen3.5-flash-02-23 | Alibaba | o1-preview, Sep 2024 | not published | Cannot tell yetIts answer was cut off mid-thought; it seemed to know events to late 2024, and its maker publishes no cutoff to compare. |
| Qwen3.6 27Bqwen/qwen3.6-27b | Chutes | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | Darkbloom | DeepSeek-R1, Jan 2025 | not published | As labelled |
| Qwen3.6 Flashqwen/qwen3.6-flash | Alibaba | DeepSeek-R1, Jan 2025 | not published | As labelled |
| Qwen3.6 Max Previewqwen/qwen3.6-max-preview | Alibaba | NYC mayoral result, Nov 2025 | not published | As labelled |
| Qwen3.6 Plusqwen/qwen3.6-plus | Alibaba | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| qwen3.6-35b-a3b by Engyengy/qwen3.6-35b-a3b | not reported | Claude 3.7 Sonnet (Feb 2025) | not published | As labelled |
| Qwen3.7 Flashqwen/qwen3.7-flash | Alibaba | DeepSeek-R1, Jan 2025 | not published | Cannot tell yetReleased Jul 2026 but showed nothing past Jan 2025; Qwen models under-report, so this is not proof of a swap. |
| Qwen3.7 Maxqwen/qwen3.7-max | Alibaba | Llama 4, Apr 2025 | not published | As labelled |
| Qwen3.7 Plusqwen/qwen3.7-plus | Alibaba | GPT-4.5 and Grok 3, Feb 2025 | not published | Cannot tell yetReleased Jun 2026 but showed nothing past Feb 2025; Qwen models under-report, so this is not proof of a swap. |
| Qwen3.8 2.4T A95Bqwen/qwen3.8-2.4t-a95b | Together | NYC mayoral result, Nov 2025 | not published | As labelled |
| Qwen3.8 27Bqwen/qwen3.8-27b | Novita | Gemini 3 launch, Nov 2025 | not published | As labelled |
| Qwen3.8 Flashqwen/qwen3.8-flash | Alibaba | no answer given | not published | Not answeringIt spent its whole answer budget thinking and gave no reply, so it could not be tested. |
| Qwen3.8 Max (0902)qwen/qwen3.8-max-0902 | not reported | Pope Leo XIV, May 2025 | not published | As labelled |
| Qwen3.8 Max Primeqwen/qwen3.8-max-prime | Alibaba | Pope Leo XIV, May 2025 | not published | As labelled |
| Qwen3.8 Omni Flashqwen/qwen3.8-omni-flash | Alibaba | Grok 4.1, Nov 2025 | not published | As labelled |
| qwen3.8-27b by Engyengy/qwen3.8-27b | not reported | GPT-5 (Aug 2025) | not published | As labelled |
| R1deepseek/deepseek-r1 | Novita | not measured | not published | ConsistentThis model gave no usable answer to our knowledge test, so we could not check it. |
| R1 0528deepseek/deepseek-r1-0528 | Novita | not measured | not published | ConsistentThis model gave no usable answer to our knowledge test, so we could not check it. |
| R1 Distill Llama 70Bdeepseek/deepseek-r1-distill-llama-70b | Novita | GPT-4o, May 2024 | Dec 2023 | As labelled |
| Reka Edgerekaai/reka-edge | Reka | nothing it could date correctly | not published | Cannot tell yetThis small model gave muddled answers, so its knowledge could not be dated. |
| Reka Flash 3rekaai/reka-flash-3 | Reka | GPT-4, Mar 2023 | not published | Cannot tell yetIts answer was cut off after one garbled item, so its knowledge could not be dated. |
| Relace Apply 3relace/relace-apply-3 | not reported | no answer | not published | Not testableA code-merge tool that only accepts code edits, so it cannot take a knowledge quiz; it needs a merge test. |
| Relace Searchrelace/relace-search | Relace | 2024 US election, Nov 2024 | not published | Not testableA code-search tool rather than a chat model, so a general-knowledge test does not apply. |
| ReMM SLERP 13Bundi95/remm-slerp-l2-13b | Mancer 2 | nothing on the ladder | Sep 2022 | ConsistentBuilt on a 2022-era base, so it cannot know anything this test asks about. |
| Sabamistralai/mistral-saba | not reported | Claude 3.7 Sonnet, Feb 2025 | not published | As labelled |
| Sakana Namazusakana/sakana-namazu | not reported | no answer | not published | Not answeringNo server could run it under our privacy settings, so we got no answer to test. |
| Schematron V2 Smallinference-net/schematron-v2-small | InferenceNet | did not say | not published | Not testableAn extraction model that turns text into JSON; it does not answer questions, so its knowledge cannot be tested. |
| Schematron V2 Turboinference-net/schematron-v2-turbo | InferenceNet | did not say | not published | Not testableAn extraction model that turns text into JSON; it does not answer questions, so its knowledge cannot be tested. |
| Seed 1.6bytedance-seed/seed-1.6 | Seed | Llama 3.1 405B (Jul 2024) | not published | As labelled |
| Seed 1.6 Flashbytedance-seed/seed-1.6-flash | Seed | Llama 3.1 (2024, month unsure) | not published | Cannot tell yetIts dates for well-known 2023 to 2024 events were mostly wrong, so its knowledge horizon could not be measured. |
| Seed 2.1 Turbobytedance-seed/seed-2-1-turbo | Seed | Claude 3.5 Sonnet update (Oct 2024) | not published | Cannot tell yetIts answers suggest knowledge ending in 2024, well before this 2026 model's release; not confirmed. |
| Seed-2.0-Codebytedance-seed/seed-2.0-code | Seed | OpenAI o1-preview (Sep 2024) | not published | Cannot tell yetIts knowledge appears to end in 2024, over a year before this model's 2026 release; the cause is unconfirmed. |
| Seed-2.0-Litebytedance-seed/seed-2.0-lite | Seed | Llama 3.1 405B (Jul 2024) | not published | Cannot tell yetIts knowledge appears to end mid-2024, well before this model's 2026 release; the cause is unconfirmed. |
| Seed-2.0-Minibytedance-seed/seed-2.0-mini | Seed | OpenAI o1-preview (Sep 2024) | not published | Cannot tell yetIts answers were inconsistent and point to knowledge ending in 2024; the cause is unconfirmed. |
| Skyfall 36B V2thedrummer/skyfall-36b-v2 | Parasail | Claude 3 family, Mar 2024 | not published | As labelled |
| Solar Mini 4upstage/solar-mini4 | Upstage | Pope Leo XIV elected, May 2025 | not published | Cannot tell yetKnew events to about May 2025, about 10 months short of what its release date suggests; not enough to call it a swap. |
| Solar Pro 3upstage/solar-pro-3 | Upstage | o1-preview, Sep 2024 | not published | Cannot tell yetRecalled events only to about late 2024 and got some wrong; served by the maker, so likely its own limit. |
| Solar Pro 4upstage/solar-pro4 | Upstage | Claude Sonnet 4, May 2025 | not published | As labelled |
| Sonarperplexity/sonar | Perplexity | Claude 3.7 Sonnet, Feb 2025 | not published | Not testableAnswers by searching the web, so a knowledge-date test does not apply. |
| Sonar Deep Researchperplexity/sonar-deep-research | Perplexity | DeepSeek-R1, Jan 2025 (via search) | not published | Not testableAnswers by searching the web, so a knowledge-date test does not apply. |
| Sonar Properplexity/sonar-pro | Perplexity | GPT-5, Aug 2025 | not published | Not testableAnswers by searching the web, so a knowledge-date test does not apply. |
| Sonar Pro Searchperplexity/sonar-pro-search | Perplexity | World Cup opener, Jun 2026 | not published | Not testableAnswers by searching the web, so a knowledge-date test does not apply. |
| Sonar Reasoning Properplexity/sonar-reasoning-pro | Perplexity | o1-preview, Sep 2024 | not published | Not testableAnswers by searching the web, so a knowledge-date test does not apply. |
| Step 3.5 Flashstepfun/step-3.5-flash | SiliconFlow | o1-preview, Sep 2024 | not published | Cannot tell yetIt showed knowledge only to late 2024; its maker publishes no cutoff, and its sibling on the maker's own service answers the same. |
| Step 3.7 Flashstepfun/step-3.7-flash | StepFun | Claude 3.5 Sonnet, Jun 2024 | not published | Cannot tell yetIt showed knowledge only to mid 2024, well short of its 2026 release; its maker's own service answers the same way. |
| Ternary Bonsai 2 27Bprism-ml/ternary-bonsai-2-27b | Darkbloom | Gemini 3, late 2025 | not published | As labelled |
| Trinity Large Thinkingarcee-ai/trinity-large-thinking | Arcee AI | DeepSeek-R1 (Jan 2025) | not published | Cannot tell yetIts knowledge seems to end around early 2025; the maker publishes no cutoff to compare against. |
| UI-TARS 7Bbytedance/ui-tars-1.5-7b | Parasail | Llama 3.1 405B (Jul 2024) | not published | As labelled |
| UnslopNemo 12Bthedrummer/unslopnemo-12b | Parasail | Llama 2, Jul 2023 | not published | As labelled |
| Venice Role Play Uncensoredvenice/venice-uncensored-role-play | not reported | Llama 3.1 405B, mid 2024 | not published | As labelled |
| Venice Uncensored 1.2venice/venice-uncensored-1-2 | not reported | Llama 3.1 405B, Jul 2024 | not published | As labelled |
| Venice: Uncensoredcognitivecomputations/dolphin-mistral-24b-venice-edition | Venice, supplier's own key | Llama 3.1 405B, Jul 2024 | Oct 2023 | As labelled |
| Voxtral Small 24B 2507mistralai/voxtral-small-24b-2507 | Mistral | OpenAI o1-preview, Sep 2024 | not published | As labelled |
| Weaver (alpha)mancer/weaver | Mancer 2 | nothing reliable | not published | ConsistentA 2023 role-play model; it invents answers about recent events, so do not rely on it for facts. |
| WizardLM-2 8x22Bmicrosoft/wizardlm-2-8x22b | Novita | OpenAI o1-preview, Sep 2024 | not published | Cannot tell yetIt knows events from after this model came out in April 2024, so a newer, different model may be answering. |
A note on the method's limits. Genuine models are cautious about the last months before their cutoff and often under-report what they know, so a short gap proves nothing and a long one is only a reason to look harder. Where a maker publishes no cutoff we use its release date minus six months. A hidden prompt can make a genuine model answer like an older one, which is why we test with the date and look at the route as well as the answers. We would rather say "cannot tell yet" than guess in either direction.
5. This is why we are already building our own router
Until this week, almost everything on askr ran through one supplier. That let us launch with hundreds of models in a few weeks. It also meant that what reached the model depended on choices made two layers below us, which we could not see and did not check. The instructions in front of Grok 4.7 are one of those choices. That is the bug behind the bug.
The fix is a router of our own: askr connects directly to the companies that make the models, one by one, and the supplier keeps only the long tail, behind the nightly check. We did not start this because of the audit. We already have one direct connection live: Claude (direct), a straight line to Anthropic's API, has been running since 25 September, and every Claude answer on it comes from Anthropic and says so. Grok comes next, directly from xAI, then GLM from Z.ai, then OpenAI, Google and DeepSeek. Each direct connection shows "(direct)" after the model's name and charges the maker's list price.
The rule from now on is short. A model on askr is the model on the label, served by its maker or by a route we name on the receipt, and checked every night.
Live
Claude (direct) to Anthropic, since 25 September. GPT (direct) to OpenAI, since 27 September. Grok (direct) to xAI, since 28 September.
Next
GLM (direct) to Z.ai.
Then
Google and DeepSeek direct. The supplier for the tail only, behind the nightly check.
6. What askr is
askr is not an inference market. We are building a consumer app, and it is already live: one place to talk to every major model, with Skills that turn a chat into a job done, Collab rooms where a group works with the models together on a shared balance, and Mind, a memory that follows you across all of it, coming soon.
The models are the ingredients. We do not make them and we do not pretend to. What we owe you is that the ingredients are what the label says, that you can see where they came from, and that when something looks wrong, you hear the whole story from us, whichever way it comes out. This page is that story. The nightly check keeps it true from now on.
7. Check it yourself
Here is how to test whether any model is the model on its label, the way we test ours now. Then the accusations made about askr, the proper test for each and what it showed, and why the original tests fell short.
How to test any model properly
- Tell it the date. Start with "Today's date is ...". Without it, a model assumes it is still early in its training and treats anything later as not yet happened. In our tests this one line turned Grok 4.7's "there was no papal election in May 2025" into "Leo XIV, elected 8 May 2025".
- Ask about the world, not about the model. Asking a model who it is proves nothing. Claude Haiku 4.5, taken straight from Anthropic, calls itself Claude 3.5 Sonnet.
- Ask for things that cannot be guessed. A winner's name, an exact release date. Skip anything announced in advance, such as where the Olympics are held or when the World Cup starts.
- Find the newest thing it gets right. Ask across two or three years of events. The newest one it gets right, with the right name and month, is where its knowledge ends. Ask the last few one at a time: in a long list, careful models say "never heard of it" too early.
- Compare with the maker's stated cutoff, and expect a gap. Genuine Claude, straight from Anthropic, knows events up to 3 to 8 months before its stated cutoff and gets vaguer after that. Only a gap of a year or more is a reason to look harder.
- Ask the maker's own copy the same way. Same date line, same questions, one at a time, with no other app in between. Official apps and coding assistants add instructions of their own, and some search the web. A difference between two like-for-like runs is the real signal.
- Make sure it is not searching the web. A model that searches knows yesterday's news and cites sources, so it tells you nothing about its training. Our table marks those "not testable".
- Count what you pay for. A bare "hi" should cost a handful of prompt tokens. Hundreds or thousands mean something sits in front of your message.
The ladder we used, short enough to paste into any chat. Add today's date in the first line.
Today's date is [today's date]. For each item, say in one line what it is and when it happened, month and year, from your own training knowledge. Try every item. If you have genuinely never heard of one, say so. 1) OpenAI's GPT-4o 2) The winner of the 2024 US presidential election 3) DeepSeek-R1 4) The pope elected in May 2025 5) OpenAI's GPT-5 6) The winner of the 2025 Nobel Peace Prize 7) Google's Gemini 3 8) The winner of the November 2025 New York City mayoral election 9) Anthropic's Claude Opus 4.6 10) The country that topped the 2026 Winter Olympics medal table 11) Google's Gemini 3.1 Pro 12) Anthropic's Claude Opus 4.8
| Item | A correct answer names |
|---|---|
| 1) OpenAI's GPT-4o | May 2024 |
| 2) The winner of the 2024 US presidential election | Donald Trump, November 2024 |
| 3) DeepSeek-R1 | January 2025 |
| 4) The pope elected in May 2025 | Leo XIV (Robert Prevost), 8 May 2025 |
| 5) OpenAI's GPT-5 | 7 August 2025 |
| 6) The winner of the 2025 Nobel Peace Prize | María Corina Machado, 10 October 2025 |
| 7) Google's Gemini 3 | 18 November 2025 |
| 8) The winner of the November 2025 New York City mayoral election | Zohran Mamdani, 4 November 2025 |
| 9) Anthropic's Claude Opus 4.6 | 5 February 2026 |
| 10) The country that topped the 2026 Winter Olympics medal table | Norway, February 2026 |
| 11) Google's Gemini 3.1 Pro | 19 February 2026 |
| 12) Anthropic's Claude Opus 4.8 | 28 May 2026 |
Norway also topped the 2018 and 2022 medal tables, so count item 10 only alongside other 2026 answers. The newest item a model gets right is where its knowledge ends; set that against its maker's stated cutoff.
The accusations, and the proper test for each
"Grok 4.7 is really an older Grok, with knowledge to mid 2025"
Did not hold up- The proper test
- Give it today's date, ask about early-2026 events one question at a time, and ask an older genuine model the same as a control.
- What we found
- It named GLM-5 (11 February 2026), Gemini 3.1 Pro (19 February 2026), Claude Opus 4.6 (5 February 2026) and Norway topping the 2026 medal table. Claude Haiku 4.5 from Anthropic, whose knowledge ends in early 2025, got the same line and could name none of them.
- Why the original fell short
- No date, so late-2025 events looked unconfirmed to the model. Most questions were about rivals' launches, which Grok models often deny: Grok 4.5 said Gemini 3.1 Pro does not exist while naming GLM-5. The official copy answered inside Cursor, which adds instructions of its own.
"Grok 4.6 is Grok 3, and its hidden script admits a 2024-10 cutoff"
Did not hold up- The proper test
- Count the prompt tokens on a bare "hi" to see whether anything sits in front of it, then ask dated questions.
- What we found
- A bare "hi" counts 19 prompt tokens: nothing of that size is in front of it. It described the 19 October 2025 Louvre theft in detail and, with the date, named María Corina Machado's Nobel Peace Prize (10 October 2025). Grok 3 came out in February 2025 and cannot know either.
- Why the original fell short
- A model asked to repeat instructions it does not have will write out the kind it was trained with. "January 2026 has not yet occurred" is exactly what any model says when nobody tells it the date.
"GLM 5.3 is really GLM-4.6"
Did not hold up- The proper test
- Ask, with the date, about events after GLM-4.6 came out on 30 September 2025, and put the same family questions to Z.ai's own model.
- What we found
- It named Gemini 3 (18 November 2025), Grok 4.1 (17 November 2025) and the 2025 Nobel Peace Prize winner. It does not know GLM-5 or GLM-4.7, and neither does GLM 5.3 FlashX served by Z.ai itself, asked the same way.
- Why the original fell short
- The family-tree questions hit a blind spot the maker's own model shares. The "real" GLM 5.3 answered inside a coding assistant on Z.ai's coding plan, which brings instructions of its own.
"It names itself three different ways"
Did not hold up- The proper test
- Do not use self-identification.
- What we found
- Genuine models do the same: Claude Haiku 4.5 from Anthropic says it is Claude 3.5 Sonnet, and Claude Opus 5 says it is Opus 4.5.
- Why the original fell short
- A model's name for itself is not evidence of anything.
"About 1,250 hidden words are billed on every Grok 4.7 message"
Right, and refunded- The proper test
- Send "hi" straight to our supplier, askr's code out of the path, and read the prompt token count.
- What we found
- 1,243 prompt tokens, 1,152 of them cached, on the Grok 4.7 route, against 19 on Grok 4.6. askr does not add them. We have refunded them and asked our supplier where they come from (section 3).
- Why the original fell short
- It held. Credit to the user.
"'The conversation is too long' comes from Grok's consumer website"
Our limit, since lifted- The proper test
- Search askr's own code and docs for the message.
- What we found
- It is askr's own error text, word for word. Until 24 September we capped a conversation at 120,000 characters; release 1.384 lifted it to each model's own context window.
- Why the original fell short
- It assumed the wording came from somewhere else. The limit was real, and it was ours.
"Its tool calls are fake"
Our gap, being fixed- The proper test
- Check whether the tool definitions reach the model at all.
- What we found
- askr's API drops them today, and the API docs say so. No model can call a tool it is never shown. The fix is in section 3.
- Why the original fell short
- It tested askr's API gap and read it as a fake model.
"About one request in four fails"
Our budget, being fixed- The proper test
- Read the finish reason on the failed calls.
- What we found
- The model spent its whole output budget thinking and wrote nothing. askr refunds those turns. A larger budget fixed nearly all of them in our tests, and the fix is in section 3.
- Why the original fell short
- A budget problem, not a model problem.
"It is shuffled between mismatched backends"
Partly right- The proper test
- Read the host that each response reports.
- What we found
- Our supplier can send the same model to a different host between requests. That changes where it runs, and it is why every receipt will name the host (section 3).
- Why the original fell short
- The hosts vary. The models on them are the ones on the label.
"Cheap models must be fake models"
Did not hold up- The proper test
- Compare what askr pays for a request with what it charges.
- What we found
- We pay our supplier its price for the named model on every request. The holder discount and the free credits come out of askr's pocket.
- Why the original fell short
- A discount tells you who pays, not which model answers.
Why the accusation tests fell short
- They never gave the model the date, so real models called real events future ones.
- They leaned on self-identification, which genuine models get wrong too.
- They compared against official models running inside other apps, with those apps' instructions, instead of like for like.
- They read askr's own limits, the message cap, the dropped tools and the thinking budget, as proof about the model.
What they got right, we have taken on: the missing date, the hidden tokens, the tools and the empty answers are all in section 3.
The short version
Paste this into any model on askr, then into the same model at its maker. Put today's date in the first line. Without it, a model that does not know what day it is will often call recent events future ones, and that is the mistake this page is about.
Today's date is [today's date]. Answer from your own training knowledge, one line each, with the month and year. Try every item. If you have genuinely never heard of one, say so. 1) The winner of the 2025 Nobel Peace Prize 2) Google's Gemini 3 3) The election of Pope Leo XIV 4) Anthropic's Claude Opus 4.6 5) The country that topped the 2026 Winter Olympics medal table 6) Exactly which model and version are you, and who made you?
Question 6 is a hint, not proof: genuine models often name an older version of themselves. A model that cannot answer 1 to 5 even with the date is worth reporting, together with the maker's stated cutoff for it.
If you find something we missed, tell us at heyaskr.ai/feedback. We credit the finder and refund anyone affected.
8. Speak to us
Everything on this page started with one person telling us something was wrong. That is how askr gets better, and we pay for it: every bug you find and every report you send earns free credits on your account. It is how we shipped 27 releases in the week since we went live on 19 September, from Collab and Skill to file attachments, web search, card payments and the direct line to Anthropic. The whole list is on heyaskr.ai/roadmap.
Feedback form
heyaskr.ai/feedback. Ten short questions, or one line about what broke. Every answer reaches the team the minute you send it.
A call with the founders
Leave your Telegram handle on the form and we book a call. We would rather hear it from you than read it in a report.
Support
Message @askrsupportbot on Telegram and you have a ticket with a person on the other end. The community is at t.me/heyaskr.
One last thing
We did not write this page to defend ourselves, and we are not here to fight anyone's accusation. A user tested us, we tested ourselves harder, and we published what we found, including the parts that were our fault.
This page exists because we care about the product. askr is a consumer app, and it gets better every time someone tells us where it breaks. We want anyone and everyone to help us build it.
Sources: the makers' model pages (xAI, Anthropic, OpenAI, Google, Z.ai and the others, as cited in the audit data), OpenRouter's public model listing, askr's API documentation, askr's own code. Test dates: 21 to 26 September 2026 (the user), 26 September 2026 (askr).