Frontier Models, September 2026: ASTRA, Fable, Jev and the Chinese Frontier
September 2026 produced the tightest frontier race on record. OpenAI's GPT-6 Astra (3 Sep) and Anthropic's Claude Fable 5.1 (1 Sep) are tied at 53 on the Artificial Analysis Intelligence Index, and they split the specialised benchmarks: Astra leads on interactive reasoning (ARC-AGI-3) and terminal agents, while Fable leads on expert knowledge (Humanity's Last Exam) and human-preference arenas. The best Chinese models (Qwen3.8 Max, GLM-5.3, Kimi K3) sit 8 to 9 index points behind. GLM-5.3 and Kimi K3 reach that score at roughly a quarter of Fable's cost, and most of the Chinese tier ships open weights. The outlier is TypeSafe's Jev (15 Sep). Jev does not generate text. It returns typed, calibrated decisions in 70 to 500 ms at $0.042 per million input tokens. We argue that Jev is not a competitor on the intelligence axis. It unbundles a workload the frontier models have been overcharging for. A sensible production stack in Q4 2026 is therefore tiered: decision models for classification and routing, open-weight Chinese or mid-tier models for bulk generation, and Fable or Astra only where depth pays for itself.
01Introduction and scope
We compare four subjects: (i) ASTRA, which we identify as OpenAI's GPT-6 Astra, not Google DeepMind's Project Astra, a research assistant prototype with no public model release[12]; (ii) FABLE, Anthropic's Claude Fable 5.1; (iii) JEV, TypeSafe AI's first "System One Model", which is unrelated to Meta's JEPA world models despite the similar acronym; and (iv) the top Chinese frontier models as of 21 September 2026. The question is practical: which model class should an operator route which work to, and at what price?
02Method
Six research agents (four Sonnet, two Haiku) searched the literature in parallel, one agent per subject plus a cross-leaderboard sweep. Every figure in this paper carries a source. Before publication, the lead author checked the load-bearing numbers against primary pages: the TypeSafe launch post, the Artificial Analysis leaderboard, the DeepSeek API changelog and TechCrunch. Each claim is labelled with its evidence tier:
- V vendor-reported (lab's own announcement or docs)
- I independent evaluator (Artificial Analysis, ARC Prize, JevBench)
- A aggregator or press (useful, lower confidence; digits may drift)
Where sources conflict, we show both values and do not average them.
03Subjects
3.1 GPT-6 Astra (OpenAI)
Launched 3 September 2026 after a limited "Daybreak" preview. The architecture uses recurrent-depth ("looped") transformers, which are more compute-efficient but make chain-of-thought less auditable. OpenAI itself flags this as a monitorability regression[1][2]. Modalities: text, vision, computer and browser use. Price: $10 / $50 per million input/output tokens, with a 90% discount on cache reads[3]. Offensive-cyber capability is strong enough that exploit development is gated behind a separate access programme[4]. Parameter count has not been disclosed.
3.2 Claude Fable 5.1 (Anthropic)
Released 1 September 2026 with a restricted sibling, Mythos 5.1. Both are the same model with different safeguard levels[5]. Context is 1M tokens, max output 128K, input is text and images. Adaptive thinking is always on. Price: $10 / $50, the same as Astra. Cache reads fell 75% to $0.25/M[6]. Anthropic's own guidance is to start with Opus 5 and move up to Fable 5.1 only for demanding reasoning or long-horizon agentic work[6].
3.3 Jev (TypeSafe AI)
Launched into early access on 15 September 2026 by Diogo Almeida (formerly OpenAI; RLHF and ChatGPT), with co-founders Erik Gafni and Sasha Sheng[7][8]. Jev is transformer-based but, by design, not a language model. It receives a block of state plus typed questions (choice, score, yes/no) and returns calibrated probabilities in a single parallel pass. It cannot emit free text. Choice cardinality is capped at 255. Training uses "Reinforcement Learning for Calibrated Decisions" (RLCD) on synthetic data, and no methodology paper has been published[7][9]. Price: $0.042 per million input tokens, output free. Latency: 70 to 500 ms[7]. Weights are closed; community reimplementations (OpenJev, SemIf, djev) copy the interface, not the model[10].
3.4 The Chinese frontier
The leading Chinese models are Qwen3.8 Max (Alibaba, GA 3 Aug, API-only), GLM-5.3 (Zhipu/Z.ai, 14 Aug, open weights under MIT in the 5.x line), Kimi K3 (Moonshot, 1M context, premium closed tier above the open K2.6), DeepSeek V4-Pro / V4.1-Flash (V4-Pro GA 13 Aug, V4.1-Flash 10 Sep with native multimodality)[13], Tencent Hunyuan Hy4 Preview (28 Aug, 770B/49B-active MoE, Apache 2.0) and StepFun Step 5 Preview (~20 Sep, 600B/27B-active)[14]. The labs are converging on one strategy: open-weight workhorses, with the newest variants behind paid APIs.
04Results
4.1 General intelligence and cost
| Model | Lab | AA Index | AA cost figure (USD) | Weights | List price in / out ($/M) |
|---|---|---|---|---|---|
| Claude Fable 5.1 (max) | Anthropic | 53 | 7.63 | Closed | 10 / 50 |
| GPT-6 Astra (max) | OpenAI | 53 | 3.26 | Closed | 10 / 50 |
| Claude Opus 5 (max) | Anthropic | 51 | 5.86 | Closed | n/a here |
| Qwen3.8 Max | Alibaba | 45 | 5.41 | Closed (API) | 2 / 6 |
| GLM-5.3 (max) | Zhipu | 45 | 2.01 | Open (MIT line) | 1.40 / 4.40 (5.2) |
| Kimi K3 (max) | Moonshot | 44 | 2.00 | Closed tier | 3 / 15 |
| DeepSeek V4.1 Flash | DeepSeek | 39 | 0.27 | Open line | 0.30 / 1.20 peak |
| MiniMax-M3 | MiniMax | 29 | 0.51 | Open line | n/a here |
| Jev | TypeSafe | not rated | n/a | Closed | 0.042 / free |
Table 1. Artificial Analysis Intelligence Index v4.x and the cost figure AA lists next to each entry, read 21 Sep 2026 I[3][11]. List prices are vendor or aggregator figures V/A[14]. Jev is out of scope for the index because it cannot produce the text answers the index grades.
4.2 Specialised benchmarks: the leaders split
| Benchmark | Fable 5.1 | GPT-6 Astra | Best Chinese | Tier |
|---|---|---|---|---|
| Humanity's Last Exam, with tools | 65.0% | 57.2% | not found | V/A [5][2] |
| Humanity's Last Exam, aggregator board | 59.1% | 54.7% | 47.6% Kimi K3 | A [15] |
| ARC-AGI-3, standard harness | not found | 62.7% | <5% | I/A [15] |
| ARC-AGI-3, OpenAI harness | n/a | 99.9% | n/a | V contested [2] |
| Terminal-Bench 4.0 (AA run) | 52% | 59% | not found | I [3] |
| Terminal-Bench 4.0 (vendor) | 55.8% | 57.7% | not found | V/A [5][2] |
| OSWorld 2.0 | 77.9% partial / 41.7% strict | 72.6% | not found | V/A scoring differs |
| AA Coding Agent Index | 62 | 62 | not found | I [3] |
| AA hallucination rate (max) | not found | 51% (from 92%) | not found | I [3] |
| LMArena text Elo | 1514 (Max) | not in top 10 | ~1460 Qwen3.7 Max | A [15] |
Table 2. Bold marks the leader in each row. "Not found" means no source was located; it does not mean the model scored zero. The OSWorld 2.0 figures use different scoring modes and are not directly comparable.
4.3 Decisioning workloads: Jev
| Measure | Jev | Frontier LLM reference | Tier |
|---|---|---|---|
| Latency per decision | 70 to 500 ms | 3 to 329 s | V [7] |
| Input price | $0.042 / M | $10 / M (Fable, Astra) | V [7][3] |
| Output price | free | $50 / M | V |
| Schema / type errors | 0% by construction | non-zero | V [7] |
| JevBench v1.2 overall (Intelligence 90.4, Calibration 82.7) | 75.4 | SemIf clone 74.7 | I [9] |
| Invoice processing accuracy (independent test) | 61.8% | 79.1% competitor | A [9] |
Table 3. Jev's claimed cost advantage is about 238 times on input tokens, and output is free. Its accuracy advantage is task-dependent and unproven. TypeSafe's own benchmark labels were partly derived from averaged Astra and Fable outputs, not ground truth. TypeSafe acknowledged this as a possible design bias[9].
05Discussion
5.1 Routing recommendation (Q4 2026)
| Workload | Default | Why |
|---|---|---|
| Classification, routing, scoring, triage (for example CX ticket intent, fraud flags) | Jev-class decision model, with an LLM fallback on low confidence | Sub-second, near-zero cost, calibrated probability usable as a threshold |
| Bulk generation and summarisation where data may leave the premises | DeepSeek V4.1 Flash / GLM-5.3 | Most capability per dollar |
| Regulated or PII data that must stay on the box | Open-weight GLM / Hunyuan self-hosted, or a US mid-tier model under a DPA | Residency. Treat Chinese-hosted APIs as out of bounds for PII |
| Long-horizon agents, expert analysis, high-stakes writing | Fable 5.1 or Astra, chosen by private eval | Depth pays; the two leaders split by task |
| Novel-environment or computer-use agents | Astra, with Fable as a second opinion | ARC-AGI-3 and Terminal-Bench lead; mind the CoT-auditability loss |
06Limitations
This is a three-week evidence window on fast-moving releases. Several benchmark digits come from aggregators, not system cards. Astra's system card was not fetchable (HTTP 403), and Jev has no architecture paper or parameter count. Chinese HLE, SWE-bench Verified and GPQA figures were mostly unavailable from primary sources. The AA cost figure is used as listed; it is AA's run-cost metric, not a per-token price. No model was run by K/20X for this paper. Every number is secondary to its cited source.
07References
- TechCrunch, "OpenAI launches Astra, its powerful (and controversial) new model", 3 Sep 2026. techcrunch.com/2026/09/03/openai-launches-astra-its-powerful-and-controversial-new-model/
- DataCamp, "GPT-6 Astra: Features, Benchmarks, and Pricing". datacamp.com/blog/gpt-6-astra
- Artificial Analysis, "Benchmarking GPT-6 Astra" and model leaderboard, read 21 Sep 2026. artificialanalysis.ai/articles/benchmarking-gpt-6-astra
- CNBC, "OpenAI begins rolling out Astra model after warning of its advanced cyber capabilities", 3 Sep 2026.
- Anthropic, "Introducing Claude Fable 5.1 and Claude Mythos 5.1", 1 Sep 2026; and System Card.
- Anthropic docs, Claude Fable 5.1 model page and models overview. platform.claude.com/docs
- TypeSafe AI, "Introducing System One Models and Jev", 15 Sep 2026. typesafe.ai/blog/introducing-system-one-models-and-jev
- TechCrunch, "A new kind of AI model from a ChatGPT inventor is thrilling developers", 18 Sep 2026.
- A. Maio, "Jev: the language model that won't" (Substack); JevBench v1.2, benchmarkheaven.com/jev-models, 21 Sep 2026.
- github.com/razorback16/openjev; github.com/cobanov/awesome-jev
- Artificial Analysis leaderboard, artificialanalysis.ai/leaderboards/models, read 21 Sep 2026.
- Google DeepMind, Project Astra. deepmind.google/models/project-astra/
- DeepSeek API change log. api-docs.deepseek.com/updates/
- DataCamp (Qwen3.8-Max); DataNorth (GLM-5.2, Seed 2.1); explainx.ai (GLM-5.3); miraflow.ai (Hunyuan Hy4); Eastern Herald (StepFun Step 5, 20 Sep 2026); BenchLM and pricepertoken.com pricing pages.
- BenchLM / llm-stats HLE and ARC-AGI-3 boards; metatext.io LMArena snapshot, 11 Sep 2026; Vercel AI Gateway leaderboard, 20 Sep 2026.