Frontier LLM Releases of Summer 2026: A Benchmark-Focused Review
A literature review of six official model releases — Claude Fable 5, Claude Opus 5, GPT-5.6, DeepSeek-V4-Flash-0731, Kimi K3, and GLM-5.2 — with a benchmark explainer and an executive summary.
Part 1 — Literature Review: What the Official Releases Claim
All findings below are taken from the official release announcements (URLs in the references section). A recurring theme in 2026: benchmarks have shifted from static knowledge tests (MMLU, GSM8K) to agentic, long-horizon evaluations — coding agents, terminal use, knowledge work, and computer use — with cost-per-task as a first-class metric alongside accuracy.
1.1 Claude Fable 5 (Anthropic, June 9, 2026; redeployed July 1)
Fable 5 is a "Mythos-class" model made safe for general use — the same underlying model as Claude Mythos 5, with heavy safety classifiers (cyber/bio/chemistry requests fall back to Opus 4.8; <5% of sessions affected). Anthropic claims it is "state-of-the-art on nearly all tested benchmarks," with the largest lead on long, complex tasks.
Benchmarks cited: FrontierCode (Cognition) — highest score among frontier models; FrontierBench — highest score; CursorBench — state of the art (per Cursor); Hebbia Finance Benchmark — highest score; ViBench (Replit) — highest. Cross-cited third-party results: GDPval-AA 1,759.6 Elo (best), AA Intelligence Index 59.9 (best), AA Coding Agent Index 77.2, SWE-bench Pro 80%, Terminal-Bench 2.1 83.1, BrowseComp 84.3%, Toolathlon 61.7, GPQA Diamond 92.6%, FrontierMath Tier 4 87.8% (best), Agents' Last Exam 40.5%, HealthBench Professional 60.9%, GraphWalks BFS 1M 79.4.
Context: Suspended June 12–30 due to US export controls following an Amazon jailbreak report; redeployed July 1 globally. Pricing \(10/\)50 per M tokens. Early testimonials emphasize months-of-engineering-in-days workloads (Stripe's 50M-line migration) and best-in-class vision (Pokémon FireRed with vision-only harness).
1.2 Claude Opus 5 (Anthropic, July 24, 2026)
Positioned as "near-Fable intelligence at half the price" and the new default on Claude Max. Anthropic claims new state-of-the-art on coding and knowledge-work evaluations.
Benchmarks cited: Frontier-Bench v0.1 — surpasses all models, more than doubles Opus 4.8's performance at lower cost-per-task; CursorBench 3.2 — within 0.5% of Fable 5 at half cost; ARC-AGI 3 — 3× the next-best model; Zapier AutomationBench — ~1.5× next-best pass rate at equal cost; OSWorld 2.0 — outperforms every model at any given cost; GDPval-AA v2 — new SOTA; HLE — best and most cost-efficient; also AutomationBench, DeepSearchQA, AA Coding Agent Index, FrontierCode 1.1 (Devin's eval), OSS-Fuzz (cyber: close to Mythos 5 at finding, far behind at exploiting), and internal life-sciences evals (spectroscopy +10.2pp over Opus 4.8, protein variant prediction +7.7pp). Priced \(5/\)25 per M tokens (same as Opus 4.8).
1.3 GPT-5.6 Sol / Terra / Luna (OpenAI, July 9, 2026 GA; July 30 price cut)
A three-tier family: Sol (flagship), Terra (balanced), Luna (cheapest). OpenAI's framing is "performance per dollar": Sol is claimed to be "state-of-the-art across coding, knowledge work, cybersecurity, and science while using fewer tokens at lower estimated cost." New ultra effort setting coordinates 4 parallel agents.
Benchmarks cited (from the release table): Agents' Last Exam 52.7% (53.6 claimed) vs Fable 5's 40.5%; AA Coding Agent Index v1.1 80 (best); Terminal-Bench 2.1 88.8 (91.9 with ultra); DeepSWE v1.1 72.7; SWE-bench Pro 64.6; BrowseComp 90.4 (92.2 ultra); OSWorld 2.0 62.6; GPQA Diamond 94.6; FrontierMath 89% (Tiers 1–3), 83% (Tier 4); MMMU Pro 84.6%; gdp.pdf 30.7%; GDPval-AA v2 1,747.8 Elo; AA Intelligence Index 58.9; HealthBench Professional 60.5%; GeneBench Pro 28.7%; LifeSciBench 59.9%; AutomationBench 18.1%; Toolathlon 58%; OpenAI MRCR v2 91.5% (256K–512K); GraphWalks BFS 79.4% (1M); ARC-AGI-3 7.78% (vs Opus 4.8's 1.5%); cybersecurity — ExploitBench 73.5% (vs GPT-5.5's 47.9%), ExploitGym 33.7%, SEC-Bench Pro 71.2%, CTF 96.7%; self-improvement — RSI Index 57.9, KernelGen 1P 61.1, NanoGPT 9.69%. Terra/Luna beat Fable 5 on Agents' Last Exam at ~1/16th the cost. July 30: Luna price cut 80%, Terra 20% (Luna now \(0.20/\)1.20 per M tokens).
1.4 DeepSeek-V4-Flash-0731 (DeepSeek, July 31, 2026)
Official release of the V4-Flash API (public beta), a re-post-trained version of V4-Flash-Preview with unchanged architecture/size. V4-Pro official release is pending. DeepSeek's headline: "significantly enhanced agent capabilities, benchmark results far exceeding V4-Pro-Preview."
Benchmarks cited: Terminal Bench 2.1 82.7; NL2Repo 54.2; Cybergym 76.7; DeepSWE 54.4; Toolathlon verified 70.3; Agent Last Exam 25.2; Automation Bench (Public) 25.1; DSBench-FullStack (internal) 68.7; DSBench-Hard (internal) 59.6. Evaluated with DeepSeek's own "Harness minimal mode" at max effort. Native Responses API support, adapted for Codex.
1.5 Kimi K3 (Moonshot AI, July 16–17, 2026)
The world's first open 2.8-trillion-parameter model (MoE, 16/896 active), native multimodal, 1M-token context, open weights released by July 27. Architecture: Kimi Delta Attention + Attention Residuals + Stable LatentMoE (~2.5× scaling efficiency vs K2). Moonshot's own verdict: "overall performance still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol," while outperforming all other tested models.
Benchmarks cited: DeepSWE v1.1 67.3 (official leaderboard, mini-SWE-agent harness); BrowseComp 90.4 (at 1M context, no context management); Terminal-Bench 2.1, Program Bench, SWE Marathon, FrontierSWE, PostTrain Bench, MLS Bench Lite, KCB 2.0 (in-house), OfficeQA Pro, SpreadsheetBench 2, MCP Atlas, AutomationBench (600-task public subset), GDPval-AA / AA-Briefcase / APEX-Agents (cited from Artificial Analysis), MMMU-Pro, PerceptionBench (Moonshot's own atomic-vision benchmark), ZeroBench. Open-ended capability demos: GPU kernel optimization (competitive with Fable 5), building a Triton-like compiler (MiniTriton) from scratch, designing a chip in a 48-hour autonomous run, and reproducing I–Love–Q astrophysics results (~2 hours vs 1–2 weeks human). Known limitations acknowledged: sensitivity to thinking-history, excessive proactiveness.
1.6 GLM-5.2 (Z.AI / Zhipu, June 13, 2026; open weights June 16)
744B-A40B MoE, 1M-token context, MIT license, built on the IndexShare sparse-attention architecture (2.9× FLOPs reduction at 1M context). Claims: "strongest open-source model on standard coding benchmarks" and a substantial leap in long-horizon capability over GLM-5.1.
Benchmarks cited (official blog; suite cross-confirmed by Kimi K3's footnotes): Terminal-Bench 2.1 81.0 (vs GLM-5.1's 63.5; within a few points of Claude Opus 4.8's 85.0, ahead of Gemini 3.1 Pro); SWE-bench Pro 62.1 (vs 58.4); plus DeepSWE, Program Bench, SWE Marathon, MLS Bench Lite, KCB 2.0, OfficeQA Pro, SpreadsheetBench 2, AutomationBench, BrowseComp, GDPval-AA. GLM-5.1 (its predecessor in the same family) held SOTA on SWE-bench Pro and NL2Repo at its release.
Part 2 — Benchmark Explainer: What Each Evaluation Tests
2.1 Software Engineering & Agentic Coding
| Benchmark | What it tests | Who runs it |
|---|---|---|
| Terminal-Bench 2.1 | Real command-line terminal usage: agents run shell commands, navigate file systems, use CLI tools to complete dev/ops tasks | TAU-bench community / AA (tbench.org) |
| SWE-bench Pro | Resolving real GitHub issues end-to-end in real codebases (harder than SWE-bench Verified; includes harder repos) | OpenAI + SWE-bench team |
| DeepSWE v1.1 | Long-horizon software engineering in large real-world codebases, sustaining work over many steps | datacurve.ai |
| Frontier-Bench v0.1 | Long-horizon engineering tasks incl. unusual tooling (e.g., rebuild a part as a FreeCAD 3D model), run on mini-SWE-agent harness; mean reward over 5 attempts | Anthropic internal (per release) |
| FrontierCode 1.1 | Hard coding tasks graded against production-quality codebase standards (code review, tests) | Cognition |
| CursorBench 3.2 | Agent performance inside the Cursor IDE workflow (edits, tests, iterations) | Cursor |
| AA Coding Agent Index v1.1 | Composite index of coding-agent performance: implementation, terminal use, real codebases | Artificial Analysis (independent) |
| NL2Repo | Generating a complete working repository from a natural-language spec | NL2Repo community |
| Program Bench | Program synthesis: generating correct code to spec across many languages | Vals AI |
| SWE Marathon | Marathon-style multi-hour engineering sessions in real repos | swe-marathon.org |
| FrontierSWE | Frontier-difficulty SWE tasks, dominance scoring | frontierswe.com |
| PostTrain Bench | Post-training/RL quality: how well models perform after reward-modeling phases | posttrainbench.com |
| MLS Bench Lite | Machine-learning systems engineering (training, tuning, infra) | MLS Bench |
| KCB 2.0 | Kimi's in-house coding benchmark (full-stack + agentic) | Moonshot AI (internal) |
| DSBench-FullStack / -Hard | DeepSeek's internal full-stack development and hard coding-agent test sets | DeepSeek (internal) |
| KernelGen 1P / NanoGPT / RSI Index | Self-improvement: kernel optimization, tiny training runs, aggregate recursive-self-improvement score | OpenAI (internal) |
| Internal Research Debugging Eval | Debugging research systems, optimizing training recipes | OpenAI (internal) |
2.2 Knowledge Work & Agentic Productivity
| Benchmark | What it tests | Who runs it |
|---|---|---|
| Agents' Last Exam | Long-running professional workflows across 55 fields, end-to-end agentic tasks | agents-last-exam.org |
| GDPval-AA | Predicting GDP per capita from economic documents — benchmark of analytical/reasoning strength | Artificial Analysis |
| AA Intelligence Index v4.1 | Broad composite: agentic work, coding, scientific reasoning, general capabilities | Artificial Analysis |
| Big Finance Bench | Financial analysis tasks (multi-hop research, quantitative reasoning) | Model ML / Rogo |
| Management Consulting Tasks | Internal consulting-style analysis workflows | OpenAI (internal) |
| AutomationBench (Zapier) | Completing business automation workflows end-to-end via tool chains | Zapier |
| OfficeQA Pro | Office/document tasks over PDF corpora rendered as images (no text layer) | Community (per Kimi) |
| SpreadsheetBench 2 | Spreadsheet formula/chart/analysis tasks | Community (per Kimi) |
| BrowseComp | Agentic web browsing: multi-step research with tool use | OpenAI |
| OSWorld 2.0 | Computer use: operating GUIs, clicking, typing, using real apps | OSWorld project |
| MCP Atlas | Tool use over MCP servers (500-task public subset, Gemini judge) | MCP community |
| Toolathlon | General tool-calling accuracy across many tools | Toolathlon project |
| DeepSearchQA | Deep multi-hop search + QA | Anthropic (internal) |
| APEX-Agents / AA-Briefcase | Agentic knowledge-work evals from Artificial Analysis | Artificial Analysis |
| Management Consulting / Finance evals (Hebbia, IMC) | Senior-level financial reasoning, document/chart interpretation | Hebbia, IMC |
2.3 Reasoning & Science
| Benchmark | What it tests | Who runs it |
|---|---|---|
| GPQA Diamond | Graduate-level, "Google-proof" multiple-choice questions in physics/chemistry/biology; experts score ~65%, expert non-specialists 34% | idavidrein/gpqa |
| FrontierMath (v2) | Original frontier mathematics problems, graded by difficulty tiers (Tier 1–3 / Tier 4) | Epoch AI |
| HLE (Humanity's Last Exam) | 2,500 expert-written frontier questions across 100+ subjects, multimodal; published in Nature | CAIS + Scale AI |
| ARC-AGI-3 | Novel abstraction/reasoning puzzles with core-knowledge priors — "easy for humans, hard for AI" | ARC Prize |
| HealthBench Professional | Professional clinical/medical reasoning | Community (per OpenAI) |
| GeneBench Pro | Long-horizon genomics & quantitative-biology analysis | OpenAI |
| LifeSciBench | Real-world biology and life-science research workflows | Community (per OpenAI) |
| MedChemBench | Medicinal-chemistry reasoning (internal) | OpenAI (internal) |
| gdp.pdf | Parsing/analyzing complex PDF documents (visual layout + content) | Artificial Analysis |
| FrontierMath Tier 4 | The hardest tier of original math research problems | Epoch AI |
2.4 Multimodal
| Benchmark | What it tests | Who runs it |
|---|---|---|
| MMMU Pro | College-level multimodal understanding (images + text) across art, science, engineering; with/without tool use | MMMU team |
| PerceptionBench | Atomic visual perception (Kimi's own benchmark for fine-grained vision) | Moonshot AI |
| ZeroBench | Zero-shot hard multimodal reasoning tasks | Community |
2.5 Long Context
| Benchmark | What it tests | Who runs it |
|---|---|---|
| OpenAI MRCR v2 | Multi-round context recall with 8 needles at 256K–512K and 512K–1M token windows | OpenAI |
| GraphWalks BFS | BFS traversal over graph structures at 256K and 1M tokens — deep long-context reasoning | Community (per OpenAI) |
2.6 Cybersecurity (frontier-tier models only)
| Benchmark | What it tests | Who runs it |
|---|---|---|
| ExploitBench | Progress from reaching vulnerable code to arbitrary code execution | OpenAI |
| ExploitGym | Turning real-world vulnerabilities into working exploits under time caps (2h/6h) | OpenAI |
| SEC-Bench Pro | Proof-of-concept generation on complex software | OpenAI/community |
| Capture-the-Flag | Solving CTF challenges (offensive security breadth) | Community |
| OSS-Fuzz | Finding then exploiting vulnerabilities in real open-source code | Anthropic (built on Google's OSS-Fuzz) |
| CyberGym | Reproducing target vulnerabilities (public leaderboard metric) | CyberGym project |
| CyScenarioBench | Success across realistic cyber scenarios | Anthropic |
Part 3 — Executive Summary: What Each Model Is Good At
Cross-model comparison on shared benchmarks (official numbers)
| Benchmark | Claude Fable 5 | Claude Opus 5 | GPT-5.6 Sol | DeepSeek V4-Flash-0731 | Kimi K3 | GLM-5.2 |
|---|---|---|---|---|---|---|
| Terminal-Bench 2.1 | 83.1 | — | 88.8 (91.9 ultra) | 82.7 | ~85 (Kimi harness) | 81.0 |
| DeepSWE v1.1 | 69.7 | — | 72.7 | 54.4 | 67.3 | ✓ (per blog) |
| SWE-bench Pro | 80.0 | — | 64.6 | — | — | 62.1 |
| BrowseComp | 84.3 | — | 90.4 (92.2 ultra) | — | 90.4 | ✓ (per blog) |
| Agents' Last Exam | 40.5 | — | 52.7 | 25.2 | — | — |
| GDPval-AA v2 (Elo) | 1,759.6 | new SOTA | 1,747.8 | — | — | ✓ (per blog) |
| AA Coding Agent Index | 77.2 | ✓ | 80 | — | — | — |
| GPQA Diamond | 92.6 | — | 94.6 | — | — | — |
| ARC-AGI-3 | — | 3× next-best | 7.78 | — | — | — |
| Toolathlon | 61.7 | — | 58.0 | 70.3 (verified) | — | — |
| OSWorld 2.0 | 54.8 (Opus 4.8) | best per-cost | 62.6 | — | — | — |
(— = not reported in the releases we reviewed; ✓ = reported in the release blog but exact figure was image-only or unextractable. Kimi K3 scores used its own harness unless noted.)
3.1 Claude Fable 5 — the all-round frontier benchmark
Best aggregate benchmark coverage of any model at release: GDPval-AA (1,759.6 Elo), AA Intelligence Index (59.9), FrontierMath Tier 4 (87.8), GraphWalks 1M long-context (79.4), HealthBench (60.9), top scores on FrontierCode, CursorBench, Hebbia Finance. Excels at vision (native multimodal, state-of-the-art per Anthropic) and very long autonomous tasks where its lead widens. Weaknesses: expensive (\(10/\)50), safety classifiers cause fallbacks, and it loses head-to-head on Agents' Last Exam (40.5 vs 52.7), ARC-AGI-3, and GPQA Diamond vs GPT-5.6 Sol.
3.2 Claude Opus 5 — the efficiency frontier at Opus tier
Not the absolute best at anything single-metric, but the best per dollar: Frontier-Bench v0.1 SOTA (2× Opus 4.8 at lower cost), ARC-AGI-3 3× next-best, OSWorld 2.0 best-at-any-cost, GDPval-AA v2 SOTA, Zapier AutomationBench ~1.5× next-best pass rate, HLE best cost-efficiency, near-Fable CursorBench (within 0.5% at half cost). Strengths: reasoning on novel problems (ARC-AGI-3), computer use, agentic coding at Opus-level pricing (\(5/\)25). Weaknesses: behind Fable on raw capability; well behind Mythos 5 on offensive cyber (exploit development).
3.3 GPT-5.6 (Sol / Terra / Luna) — the strongest coding & efficiency workhorse
Sol is the benchmark king on agentic coding (AA Coding Index 80, Terminal-Bench 88.8, DeepSWE 72.7), browsing (BrowseComp 92.2 ultra), computer use (OSWorld 62.6), academics (GPQA 94.6, FrontierMath 89), long context (MRCR 91.5), cybersecurity (ExploitBench 73.5, CTF 96.7, SEC-Bench Pro 71.2), and professional knowledge work (Agents' Last Exam 53.6 — 13 points over Fable 5). Terra/Luna extend the family's price-performance: Luna outperforms Fable 5 on Agents' Last Exam at ~1/16th the cost, and matches year-old frontier models at ~6% of the cost. Weaknesses: ARC-AGI-3 is low (7.78%) — novel abstraction is its blind spot; AutomationBench (18.1) is the worst of the frontier group; ultra's multi-agent parallelism is the only way to its best scores.
3.4 DeepSeek-V4-Flash-0731 — the agentic open-source cost leader
A Flash-class model that leads on terminal/agentic tool use among open models: Terminal-Bench 2.1 (82.7, near GPT-5.6 Sol), Toolathlon verified (70.3 — the top reported figure here), Cybergym (76.7), NL2Repo (54.2), Automation Bench Public (25.1). Strengths: very strong agent harness integration (Responses API, Codex-adapted), low cost. Weaknesses: far behind the frontier on professional knowledge work (Agent Last Exam 25.2 vs 52.7) and full-stack engineering depth (DeepSWE 54.4); it's a Flash-tier model — V4-Pro (still pending) is the intended frontier competitor.
3.5 Kimi K3 — the best open-weight long-horizon model
The first open 3T-class model. Strongest claims on long-horizon agentic work: BrowseComp 90.4 (tied with GPT-5.6 Sol, with 1M context and no context management), DeepSWE 67.3, plus leadership on in-house productivity suites (OfficeQA Pro, SpreadsheetBench 2, MCP Atlas) and open-ended capability (kernel optimization competitive with Fable 5; built a working Triton-like compiler; designed a chip autonomously). Native multimodal + 1M context + open weights (by Jul 27). Weaknesses: officially concedes a UX/capability gap vs Claude Fable 5 and GPT-5.6 Sol; sensitivity to thinking-history in non-Kimi harnesses; excessive proactiveness on ambiguous tasks.
3.6 GLM-5.2 — the strongest open-source coding model
Best open-source performance on standard coding benchmarks: Terminal-Bench 2.1 (81.0 — within a few points of Claude Opus 4.8's 85.0, ahead of Gemini 3.1 Pro) and SWE-bench Pro (62.1, beating GPT-5.6's mid-tier results from the open-source seat). MIT license, 1M context, IndexShare architecture at 2.9× FLOPs savings. Weaknesses: as an open model it trails the closed frontier on aggregate intelligence (AA Index, Agents' Last Exam class evals); its edge is concentrated in coding/long-horizon engineering rather than breadth.
Bottom line
- Coding/agents: GPT-5.6 Sol (closed), GLM-5.2 & Kimi K3 (open)
- Novel reasoning (ARC-AGI-3): Claude Opus 5
- Knowledge work/analytics: Claude Fable 5 & GPT-5.6 Sol (tie, different benchmarks), GPT-5.6 Luna (cost)
- Long-context + multimodal: Claude Fable 5, Kimi K3 (open), GPT-5.6 Sol
- Cyber: Claude Mythos 5 (restricted), GPT-5.6 Sol (generally available)
- Price-performance: Claude Opus 5 (premium tier), GPT-5.6 Luna/Terra, DeepSeek-V4-Flash-0731 (open)
References
- Claude Fable 5 / Mythos 5 launch: https://www.anthropic.com/news/claude-fable-5-mythos-5 · redeployment: https://www.anthropic.com/news/redeploying-fable-5
- Claude Opus 5: https://www.anthropic.com/news/claude-opus-5
- GPT-5.6 launch: https://openai.com/index/gpt-5-6/ · price/performance update: https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/
- DeepSeek-V4-Flash-0731 (change log): https://api-docs.deepseek.com/updates · API models: https://api-docs.deepseek.com/
- Kimi K3 tech blog: https://www.kimi.com/blog/kimi-k3 · weights: https://github.com/MoonshotAI/Kimi-K3 · API: https://platform.kimi.ai/
- GLM-5.2 release: https://z.ai/blog/glm-5.2 · GitHub: https://github.com/zai-org/GLM-5 · model: https://huggingface.co/zai-org/GLM-5.2
- Benchmark origin sites referenced in releases: https://www.swebench.com/ · https://www.frontierbench.ai/ · https://agents-last-exam.org/ · https://artificialanalysis.ai/evaluations/gdpval-aa · https://artificialanalysis.ai/ · https://deepswe.datacurve.ai/ · https://www.swe-marathon.org/ · https://www.frontierswe.com/ · https://posttrainbench.com/ · https://www.vals.ai/benchmarks/programbench · https://lastexam.ai/ · https://arcprize.org/arc-agi · https://livecodebench.github.io/