Skip to content

LLM Landscape — Latest Snapshot

A rolling digest, current as of early August 2026. Prices and benchmark values change quickly; the linked primary sources and dated snapshot pages are the authority for current values. Full data tables live in the snapshot archive below — this page is the quick orientation.


1. Landscape at a glance

Six releases dominated summer 2026. See the frontier model benchmark review for the full literature review and benchmark explainer.

Tier Models One-line take
Frontier (closed) Claude Fable 5, Claude Opus 5, GPT-5.6 Sol Best aggregate benchmarks (Fable 5), best agentic coding (GPT-5.6 Sol), best per-dollar at premium tier (Opus 5)
Workhorse (closed) GPT-5.6 Terra/Luna, Claude Sonnet 5 Price-performance: Luna got an 80% price cut on Jul 30
Frontier (open weights) Kimi K3, GLM-5.2, DeepSeek V4 Flash 0731 Kimi K3 leads open long-horizon agentic work; GLM-5.2 is the strongest open coding model; V4 Flash 0731 is the agentic cost leader

Intelligence Index (Artificial Analysis, max-effort / adaptive reasoning):

Rank Model Index
1 Kimi K3 (max) 57
2 Claude Opus 4.8 56
3 GPT-5.6 Terra (max) 55
4 Claude Sonnet 5 53
5 GPT-5.6 Luna (max) 51
7 GLM-5.2 (max) 51
8 DeepSeek V4 Flash 0731 50
10 DeepSeek V4 Pro 44

Note: Claude Fable 5 / GPT-5.6 Sol / Claude Opus 5 score higher (AA Intelligence Index ~58–60) but are evaluated on newer index versions; treat this table as the current-generation ranking where available.

Quick picks: - Coding / agents: GPT-5.6 Sol (closed), GLM-5.2 & Kimi K3 (open) - Novel reasoning (ARC-AGI-3): Claude Opus 5 - Knowledge work: Claude Fable 5, GPT-5.6 Sol - Long context + multimodal: Claude Fable 5, Kimi K3, GPT-5.6 Sol - Price-performance: Claude Opus 5 (premium), GPT-5.6 Luna/Terra, DeepSeek V4 Flash 0731 (open)


2. Pricing at a glance

All prices USD per 1M tokens, cache-miss input / output, list price.

Frontier closed models

Model Input Output Context
Claude Fable 5 $10.00 $50.00 1M
Claude Opus 5 $5.00 $25.00 1M
GPT-5.6 Sol $5.00 $30.00 256K
GPT-5.6 Terra $2.50 $15.00 256K
GPT-5.6 Luna $0.20 $1.20 256K (after Jul 30 cut)
Claude Sonnet 5 \(2.00–\)3.00 \(10.00–\)15.00 1M

Open-weight / cheap

Model Input Output Context
DeepSeek V4 Flash 0731 $0.14 $0.28 1M (cache-hit input $0.0028)
DeepSeek V4 Pro $0.435 $0.87 1M
GLM-5.2 ~$1.12 ~$3.92 1M
Kimi K3 ~$2.80 ~$14.00 1M
Muse Spark 1.1 $1.25 $4.25

Reference task cost (20,000 in + 5,000 out tokens, no cache)

Model Task cost
DeepSeek V4 Flash 0731 $0.0042
GPT-5.6 Luna $0.0500
Gemini 3.6 Flash $0.0675
Claude Sonnet 5 $0.1350
Claude Opus 4.8 $0.2250

These are arithmetic estimates for a fixed token mix, not observed task costs. See Model intelligence & cost per task and the DeepSeek V4 Flash 0731 snapshot for the methodology and caveats.


3. Snapshot archive

Dated research snapshots with the full tables behind the digest above:

Snapshot Date What it contains
LLM Landscape Snapshot (full) 2026-07-19 Full model list, Text/Agent Arena leaderboards, per-provider API pricing tables, value analysis
LLM Landscape Snapshot (week update) 2026-07-26 New releases (Claude Opus 5, Gemini Flash Cyber), product launches, incidents, funding
Chinese LLM API Pricing 2026-07-26 Token pricing for DeepSeek, Kimi K3, GLM-5.2, Grok 4.5, GPT-5.6, Gemini 3.6 Flash, Muse Spark 1.1
DeepSeek V4 Flash 0731 — ranking & cost 2026-08-01 DeepSeek's efficiency-tier release: Intelligence Index, agent benchmarks, cost-per-task analysis
Model intelligence & cost per task 2026-07-26 Benchmark score vs token use vs task cost methodology and comparison