LLM Landscape — Latest Snapshot

Rolling digest of the LLM landscape: current frontier models, pricing at a glance, and links to the full dated snapshots.

A rolling digest, current as of early August 2026. Prices and benchmark values change quickly; the linked primary sources and dated snapshot pages are the authority for current values. Full data tables live in the snapshot archive below — this page is the quick orientation.


1. Landscape at a glance

Six releases dominated summer 2026. See the frontier model benchmark review for the full literature review and benchmark explainer.

TierModelsOne-line take
Frontier (closed)Claude Fable 5, Claude Opus 5, GPT-5.6 SolBest aggregate benchmarks (Fable 5), best agentic coding (GPT-5.6 Sol), best per-dollar at premium tier (Opus 5)
Workhorse (closed)GPT-5.6 Terra/Luna, Claude Sonnet 5Price-performance: Luna got an 80% price cut on Jul 30
Frontier (open weights)Kimi K3, GLM-5.2, DeepSeek V4 Flash 0731Kimi K3 leads open long-horizon agentic work; GLM-5.2 is the strongest open coding model; V4 Flash 0731 is the agentic cost leader

Intelligence Index (Artificial Analysis, max-effort / adaptive reasoning):

RankModelIndex
1Kimi K3 (max)57
2Claude Opus 4.856
3GPT-5.6 Terra (max)55
4Claude Sonnet 553
5GPT-5.6 Luna (max)51
7GLM-5.2 (max)51
8DeepSeek V4 Flash 073150
10DeepSeek V4 Pro44

Note: Claude Fable 5 / GPT-5.6 Sol / Claude Opus 5 score higher (AA Intelligence Index ~58–60) but are evaluated on newer index versions; treat this table as the current-generation ranking where available.

Quick picks:

  • Coding / agents: GPT-5.6 Sol (closed), GLM-5.2 & Kimi K3 (open)
  • Novel reasoning (ARC-AGI-3): Claude Opus 5
  • Knowledge work: Claude Fable 5, GPT-5.6 Sol
  • Long context + multimodal: Claude Fable 5, Kimi K3, GPT-5.6 Sol
  • Price-performance: Claude Opus 5 (premium), GPT-5.6 Luna/Terra, DeepSeek V4 Flash 0731 (open)

2. Pricing at a glance

All prices USD per 1M tokens, cache-miss input / output, list price.

Frontier closed models

ModelInputOutputContext
Claude Fable 5$10.00$50.001M
Claude Opus 5$5.00$25.001M
GPT-5.6 Sol$5.00$30.00256K
GPT-5.6 Terra$2.50$15.00256K
GPT-5.6 Luna$0.20$1.20256K (after Jul 30 cut)
Claude Sonnet 5$2.00–$3.00$10.00–$15.001M

Open-weight / cheap

ModelInputOutputContext
DeepSeek V4 Flash 0731$0.14$0.281M (cache-hit input $0.0028)
DeepSeek V4 Pro$0.435$0.871M
GLM-5.2~$1.12~$3.921M
Kimi K3~$2.80~$14.001M
Muse Spark 1.1$1.25$4.25—

Reference task cost (20,000 in + 5,000 out tokens, no cache)

ModelTask cost
DeepSeek V4 Flash 0731$0.0042
GPT-5.6 Luna$0.0500
Gemini 3.6 Flash$0.0675
Claude Sonnet 5$0.1350
Claude Opus 4.8$0.2250

These are arithmetic estimates for a fixed token mix, not observed task costs. See Model intelligence & cost per task and the DeepSeek V4 Flash 0731 snapshot for the methodology and caveats.


3. Snapshot archive

Dated research snapshots with the full tables behind the digest above:

SnapshotDateWhat it contains
LLM Landscape Snapshot (full)2026-07-19Full model list, Text/Agent Arena leaderboards, per-provider API pricing tables, value analysis
LLM Landscape Snapshot (week update)2026-07-26New releases (Claude Opus 5, Gemini Flash Cyber), product launches, incidents, funding
Chinese LLM API Pricing2026-07-26Token pricing for DeepSeek, Kimi K3, GLM-5.2, Grok 4.5, GPT-5.6, Gemini 3.6 Flash, Muse Spark 1.1
DeepSeek V4 Flash 0731 — ranking & cost2026-08-01DeepSeek’s efficiency-tier release: Intelligence Index, agent benchmarks, cost-per-task analysis
Model intelligence & cost per task2026-07-26Benchmark score vs token use vs task cost methodology and comparison