--- title: "Frontier Models, September 2026: ASTRA, Fable, Jev and the Chinese Frontier" author: K/20X LABS date: 2026-09-21 canonical: https://k20x.com/research/frontier-models-2026-09/ license: Quotable with attribution and a link --- # Frontier Models, September 2026: ASTRA, Fable, Jev and the Chinese Frontier K/20X LABS research desk · Published 21 September 2026 · Evidence window 1 to 21 September 2026 · ~12 min read **ABSTRACT** September 2026 produced the tightest frontier race on record. OpenAI's **GPT-6 Astra** (3 Sep) and Anthropic's **Claude Fable 5.1** (1 Sep) are tied at 53 on the Artificial Analysis Intelligence Index, and they split the specialised benchmarks: Astra leads on interactive reasoning (ARC-AGI-3) and terminal agents, while Fable leads on expert knowledge (Humanity's Last Exam) and human-preference arenas. The best Chinese models (Qwen3.8 Max, GLM-5.3, Kimi K3) sit 8 to 9 index points behind. GLM-5.3 and Kimi K3 reach that score at roughly a quarter of Fable's cost, and most of the Chinese tier ships open weights. The outlier is TypeSafe's **Jev** (15 Sep). Jev does not generate text. It returns typed, calibrated decisions in 70 to 500 ms at $0.042 per million input tokens. We argue that Jev is not a competitor on the intelligence axis. It unbundles a workload the frontier models have been overcharging for. A sensible production stack in Q4 2026 is therefore tiered: decision models for classification and routing, open-weight Chinese or mid-tier models for bulk generation, and Fable or Astra only where depth pays for itself. ## 01. Introduction and scope We compare four subjects: (i) **ASTRA**, which we identify as OpenAI's GPT-6 Astra, not Google DeepMind's Project Astra, a research assistant prototype with no public model release[12]; (ii) **FABLE**, Anthropic's Claude Fable 5.1; (iii) **JEV**, TypeSafe AI's first "System One Model", which is unrelated to Meta's JEPA world models despite the similar acronym; and (iv) the top Chinese frontier models as of 21 September 2026. The question is practical: which model class should an operator route which work to, and at what price? ## 02. Method Six research agents (four Sonnet, two Haiku) searched the literature in parallel, one agent per subject plus a cross-leaderboard sweep. Every figure in this paper carries a source. Before publication, the lead author checked the load-bearing numbers against primary pages: the TypeSafe launch post, the Artificial Analysis leaderboard, the DeepSeek API changelog and TechCrunch. Each claim is labelled with its evidence tier: - V vendor-reported (lab's own announcement or docs) - I independent evaluator (Artificial Analysis, ARC Prize, JevBench) - A aggregator or press (useful, lower confidence; digits may drift) Where sources conflict, we show both values and do not average them. ## 03. Subjects ### 3.1 GPT-6 Astra (OpenAI) Launched 3 September 2026 after a limited "Daybreak" preview. The architecture uses recurrent-depth ("looped") transformers, which are more compute-efficient but make chain-of-thought less auditable. OpenAI itself flags this as a monitorability regression[1][2]. Modalities: text, vision, computer and browser use. Price: **$10 / $50** per million input/output tokens, with a 90% discount on cache reads[3]. Offensive-cyber capability is strong enough that exploit development is gated behind a separate access programme[4]. Parameter count has not been disclosed. ### 3.2 Claude Fable 5.1 (Anthropic) Released 1 September 2026 with a restricted sibling, Mythos 5.1. Both are the same model with different safeguard levels[5]. Context is 1M tokens, max output 128K, input is text and images. Adaptive thinking is always on. Price: **$10 / $50**, the same as Astra. Cache reads fell 75% to $0.25/M[6]. Anthropic's own guidance is to start with Opus 5 and move up to Fable 5.1 only for demanding reasoning or long-horizon agentic work[6]. ### 3.3 Jev (TypeSafe AI) Launched into early access on 15 September 2026 by Diogo Almeida (formerly OpenAI; RLHF and ChatGPT), with co-founders Erik Gafni and Sasha Sheng[7][8]. Jev is transformer-based but, by design, *not* a language model. It receives a block of state plus typed questions (choice, score, yes/no) and returns calibrated probabilities in a single parallel pass. It cannot emit free text. Choice cardinality is capped at 255. Training uses "Reinforcement Learning for Calibrated Decisions" (RLCD) on synthetic data, and no methodology paper has been published[7][9]. Price: **$0.042 per million input tokens, output free**. Latency: 70 to 500 ms[7]. Weights are closed; community reimplementations (OpenJev, SemIf, djev) copy the interface, not the model[10]. ### 3.4 The Chinese frontier The leading Chinese models are **Qwen3.8 Max** (Alibaba, GA 3 Aug, API-only), **GLM-5.3** (Zhipu/Z.ai, 14 Aug, open weights under MIT in the 5.x line), **Kimi K3** (Moonshot, 1M context, premium closed tier above the open K2.6), **DeepSeek V4-Pro / V4.1-Flash** (V4-Pro GA 13 Aug, V4.1-Flash 10 Sep with native multimodality)[13], **Tencent Hunyuan Hy4 Preview** (28 Aug, 770B/49B-active MoE, Apache 2.0) and **StepFun Step 5 Preview** (~20 Sep, 600B/27B-active)[14]. The labs are converging on one strategy: open-weight workhorses, with the newest variants behind paid APIs. ## 04. Results ### 4.1 General intelligence and cost | Model | Lab | AA Index | AA cost figure (USD) | Weights | List price in / out ($/M) | |---|---|---|---|---|---| | Claude Fable 5.1 (max) | Anthropic | 53 | 7.63 | Closed | 10 / 50 | | GPT-6 Astra (max) | OpenAI | 53 | 3.26 | Closed | 10 / 50 | | Claude Opus 5 (max) | Anthropic | 51 | 5.86 | Closed | n/a here | | Qwen3.8 Max | Alibaba | 45 | 5.41 | Closed (API) | 2 / 6 | | GLM-5.3 (max) | Zhipu | 45 | 2.01 | Open (MIT line) | 1.40 / 4.40 (5.2) | | Kimi K3 (max) | Moonshot | 44 | 2.00 | Closed tier | 3 / 15 | | DeepSeek V4.1 Flash | DeepSeek | 39 | 0.27 | Open line | 0.30 / 1.20 peak | | MiniMax-M3 | MiniMax | 29 | 0.51 | Open line | n/a here | | Jev | TypeSafe | not rated | n/a | Closed | 0.042 / free | Table 1. Artificial Analysis Intelligence Index v4.x and the cost figure AA lists next to each entry, read 21 Sep 2026 I[3][11]. List prices are vendor or aggregator figures V/A[14]. Jev is out of scope for the index because it cannot produce the text answers the index grades. *Figure 1. Capability vs cost, Artificial Analysis data read 21 Sep 2026. The US labs lead by about 8 points. At the same score, Astra costs less than half of what Fable costs. DeepSeek V4.1 Flash is the value point: 74% of Fable's index score for 3.5% of its cost figure.* ### 4.2 Specialised benchmarks: the leaders split | Benchmark | Fable 5.1 | GPT-6 Astra | Best Chinese | Tier | |---|---|---|---|---| | Humanity's Last Exam, with tools | 65.0% | 57.2% | not found | V/A [5][2] | | Humanity's Last Exam, aggregator board | 59.1% | 54.7% | 47.6% Kimi K3 | A [15] | | ARC-AGI-3, standard harness | not found | 62.7% | <5% | I/A [15] | | ARC-AGI-3, OpenAI harness | n/a | 99.9% | n/a | V contested [2] | | Terminal-Bench 4.0 (AA run) | 52% | 59% | not found | I [3] | | Terminal-Bench 4.0 (vendor) | 55.8% | 57.7% | not found | V/A [5][2] | | OSWorld 2.0 | 77.9% partial / 41.7% strict | 72.6% | not found | V/A scoring differs | | AA Coding Agent Index | 62 | 62 | not found | I [3] | | AA hallucination rate (max) | not found | 51% (from 92%) | not found | I [3] | | LMArena text Elo | 1514 (Max) | not in top 10 | ~1460 Qwen3.7 Max | A [15] | Table 2. Bold marks the leader in each row. "Not found" means no source was located; it does not mean the model scored zero. The OSWorld 2.0 figures use different scoring modes and are not directly comparable. ### 4.3 Decisioning workloads: Jev | Measure | Jev | Frontier LLM reference | Tier | |---|---|---|---| | Latency per decision | 70 to 500 ms | 3 to 329 s | V [7] | | Input price | $0.042 / M | $10 / M (Fable, Astra) | V [7][3] | | Output price | free | $50 / M | V | | Schema / type errors | 0% by construction | non-zero | V [7] | | JevBench v1.2 overall (Intelligence 90.4, Calibration 82.7) | 75.4 | SemIf clone 74.7 | I [9] | | Invoice processing accuracy (independent test) | 61.8% | 79.1% competitor | A [9] | Table 3. Jev's claimed cost advantage is about 238 times on input tokens, and output is free. Its accuracy advantage is task-dependent and unproven. TypeSafe's own benchmark labels were partly derived from averaged Astra and Fable outputs, not ground truth. TypeSafe acknowledged this as a possible design bias[9]. ## 05. Discussion **F1. The top is a tie, and the tie breaks on workload.** Fable 5.1 and Astra share the AA headline. Astra wins on interactive or novel-environment reasoning (ARC-AGI-3) and terminal agents. Fable wins on expert knowledge (HLE) and on human preference. At equal list prices, Astra's AA cost figure is 57% lower, so for bulk agentic runs Astra is the cheaper route to the same index score. **F2. China trails by about one generation on capability but not on cost.** The best Chinese entries score 44 to 45, which is roughly where the US frontier stood one to two releases ago (GPT-5.6 Sol scores 47). GLM-5.3 and Kimi K3 reach that score at about a quarter of Fable's cost, and DeepSeek V4.1 Flash reaches 39 at 3.5% of it. The strategic moat is openness. MIT and Apache weights (GLM, Hunyuan, the StepFun and MiniMax lines) can run air-gapped, which no US frontier model allows. **F3. Jev is an unbundling, not a rival.** Most production "AI" calls in operator stacks are decisions: route this ticket, score this lead, is this a fraud pattern, which segment. Paying $50/M output tokens and waiting seconds for a reasoning model to write "yes" is the inefficiency Jev targets. Jev's fast adoption (reported #1 on Vercel's AI Gateway share board, 20 Sep[15]) signals demand for that unbundling. The claim that it "cannot hallucinate" is about output shape, not correctness. A well-formed wrong probability is still wrong[9]. **F4. Evaluation itself is now the contested ground.** Astra's ARC-AGI-3 score moves 37 points depending on the harness. Fable's OSWorld moves 36 points depending on the scoring mode. Aggregators disagree on HLE by up to 6 points. Any buying decision should rest on a private eval run on your own tasks, not on a headline number. ### 5.1 Routing recommendation (Q4 2026) | Workload | Default | Why | |---|---|---| | Classification, routing, scoring, triage (for example CX ticket intent, fraud flags) | Jev-class decision model, with an LLM fallback on low confidence | Sub-second, near-zero cost, calibrated probability usable as a threshold | | Bulk generation and summarisation where data may leave the premises | DeepSeek V4.1 Flash / GLM-5.3 | Most capability per dollar | | Regulated or PII data that must stay on the box | Open-weight GLM / Hunyuan self-hosted, or a US mid-tier model under a DPA | Residency. Treat Chinese-hosted APIs as out of bounds for PII | | Long-horizon agents, expert analysis, high-stakes writing | Fable 5.1 or Astra, chosen by private eval | Depth pays; the two leaders split by task | | Novel-environment or computer-use agents | Astra, with Fable as a second opinion | ARC-AGI-3 and Terminal-Bench lead; mind the CoT-auditability loss | ## 06. Limitations This is a three-week evidence window on fast-moving releases. Several benchmark digits come from aggregators, not system cards. Astra's system card was not fetchable (HTTP 403), and Jev has no architecture paper or parameter count. Chinese HLE, SWE-bench Verified and GPQA figures were mostly unavailable from primary sources. The AA cost figure is used as listed; it is AA's run-cost metric, not a per-token price. No model was run by K/20X for this paper. Every number is secondary to its cited source. ## 07. References [1] TechCrunch, "OpenAI launches Astra, its powerful (and controversial) new model", 3 Sep 2026. techcrunch.com/2026/09/03/openai-launches-astra-its-powerful-and-controversial-new-model/ [2] DataCamp, "GPT-6 Astra: Features, Benchmarks, and Pricing". datacamp.com/blog/gpt-6-astra [3] Artificial Analysis, "Benchmarking GPT-6 Astra" and model leaderboard, read 21 Sep 2026. artificialanalysis.ai/articles/benchmarking-gpt-6-astra [4] CNBC, "OpenAI begins rolling out Astra model after warning of its advanced cyber capabilities", 3 Sep 2026. [5] Anthropic, "Introducing Claude Fable 5.1 and Claude Mythos 5.1", 1 Sep 2026; and System Card. [6] Anthropic docs, Claude Fable 5.1 model page and models overview. platform.claude.com/docs [7] TypeSafe AI, "Introducing System One Models and Jev", 15 Sep 2026. typesafe.ai/blog/introducing-system-one-models-and-jev [8] TechCrunch, "A new kind of AI model from a ChatGPT inventor is thrilling developers", 18 Sep 2026. [9] A. Maio, "Jev: the language model that won't" (Substack); JevBench v1.2, benchmarkheaven.com/jev-models, 21 Sep 2026. [10] github.com/razorback16/openjev; github.com/cobanov/awesome-jev [11] Artificial Analysis leaderboard, artificialanalysis.ai/leaderboards/models, read 21 Sep 2026. [12] Google DeepMind, Project Astra. deepmind.google/models/project-astra/ [13] DeepSeek API change log. api-docs.deepseek.com/updates/ [14] DataCamp (Qwen3.8-Max); DataNorth (GLM-5.2, Seed 2.1); explainx.ai (GLM-5.3); miraflow.ai (Hunyuan Hy4); Eastern Herald (StepFun Step 5, 20 Sep 2026); BenchLM and pricepertoken.com pricing pages. [15] BenchLM / llm-stats HLE and ARC-AGI-3 boards; metatext.io LMArena snapshot, 11 Sep 2026; Vercel AI Gateway leaderboard, 20 Sep 2026. Canonical HTML: https://k20x.com/research/frontier-models-2026-09/