Benchmark Cost Frontiers — Top-10 Performance-per-Dollar Across Coding and Non-Coding Boards (2026-09-20)

A dated snapshot ranking model configurations by performance per dollar on five boards (DeepSWE v1.1, Terminal-Bench 4.0, ARC-AGI-2, SWE-bench Pro, Vending-Bench 2), with published costs listed per row, the DeepSWE Pareto frontier computed two ways, and a coverage check for GLM-5.3-Flash and DeepSeek. Rerun procedure lives in templates/research/cost-frontier-benchmarks-rerun.md.

Snapshot: data fetched live on 2026-09-17 (DeepSWE board last updated by its operator on 2026-09-03; other boards as fetched the same day). Costs are the boards’ own published figures on the day each run happened — read them as dated receipts, not current quotes. Rerun checklist: templates/research/cost-frontier-benchmarks-rerun.md in this repo.

What “top 10 by performance/cost” means here

Three readings exist, and they disagree. This snapshot separates them instead of blending:

  1. Ratio rankingscore ÷ published cost per task. What the tables below show. Simple and reproducible, but the ratio has known flaws (see caveats): it rewards low scorers on cheap boards, and its level sets are hyperbolas, not a preference.
  2. Pareto frontier — the non-dominated set: a config stays only if no cheaper-or-equal config has an equal-or-better score. This is the preference-free answer to “which models should exist on the menu.”
  3. Cost per solved taskcost ÷ (score × tasks). A different lens that reorders boards; it flatters models that attempt cheaply and fail cheaply.

All costs are the benchmark operators’ published figures (vendor-run for their own benchmarks; independent operators for cross-model boards). Where a board does not publish per-model cost, the table says so rather than inventing one.

DeepSWE v1.1 (agentic coding, 113 tasks) — cost per task published

Board: deepswe.datacurve.ai, 21 “best” configurations, updated 2026-09-03. Metric: Pass@1 ± CI under a fixed mini-swe-agent harness. Top 10 of 21 by Pass@1 per dollar:

Model (effort)Pass@1Avg cost/taskPass@1 per $
GLM-5.3-flash (max)63%±4%$0.24262.5
DeepSeek V4 Flash (max)53%±4%$0.46115.2
GPT-5.6 Luna (max)67%±4%$0.61109.8
DeepSeek V4 Pro (max)63%±6%$1.6737.7
Gemini 3.7 Flash (medium)65%±3%$2.0332.0
Gemini 3.8 Flash (high)74%±1%$2.3631.4
Gemini 3.6 Flash (high)47%±4%$2.2121.3
Grok 4.6 (medium)67%±2%$3.4519.4
GLM-5.3 (max)69%±3%$3.9917.3
Qwen3.8-max (xhigh)57%±3%$3.7315.3

The same board under the strict Pareto reading keeps only three configs:

FrontierPass@1Cost/task
GLM-5.3-flash (max)63%±4%$0.24
GPT-5.6 Luna (max)67%±4%$0.61
Gemini 3.8 Flash (high)74%±1%$2.36

Everything else is dominated — including the expensive tier: Claude Opus 5 (74% at $11.84) and GPT-6 Astra (74% at $6.52) are deleted by Gemini 3.8 Flash (74% at $2.36, equal score, lower cost); Claude Fable 5 (70% at $13.41) by the same point. Arithmetic (cost ascending, keep strictly better running-max score) was cross-checked against an independent O(n²) dominance test — identical set. Within the ±4-point confidence intervals, GLM-5.3-flash vs GPT-5.6 Luna overlap, so “dominated” here means point estimates, not proven gaps.

Terminal-Bench 4.0 (agentic terminal, 66 tasks) — total run cost published

Board: tbench.ai. Publishes resolution rate (±95% CI) and total COST and TOKENS per run; the ratio below divides by the published whole-board run cost, so it is a derived ranking. Rows use different agent harnesses (Codex, Claude Code, Grok Build, mini-SWE-agent), which the board itself flags per row — treat harness differences as part of the measurement. 14 rows exist; top 10:

Model (effort)AgentResolutionRun costRes.% per $1k
GPT-5.6 Luna (max)Codex17.3%±2.8%$0.3k57.7
GPT-6 Astra (max)Codex58.2%±2.8%$3.3k17.6
GLM-5.3 (max)Claude Code41.8%±3.2%$2.7k15.5
GPT-5.6 Sol (max)Codex37.3%±3.8%$2.5k14.9
GPT-5.6 Terra (max)Codex21.5%±3.3%$1.7k12.6
Gemini 3.8 Flash (high)mini-SWE-agent19.1%±3.4%$1.8k10.6
Fable 5.1 (max)Claude Code57.9%±3.8%$6.2k9.3
Opus 5 (xhigh)Claude Code53.9%±3.2%$6.1k8.8
Gemini 3.7 Flash (high)mini-SWE-agent11.2%±2.4%$1.3k8.6
Fable 5 (max)Claude Code44.5%±3.8%$7.3k6.1

Notes: no GLM-5.3-flash and no DeepSeek row on the current board (GLM-5.3 is the family’s representative). GPT-5.6 Luna’s ratio is inflated by a very low score — ratio rewards doing a cheap subset (the same trap TokenCost documented for cost-per-solved-task on TB 3.0). Attempts-per-task for TB 4.0 runs are not published on the board I could fetch, so absolute per-task costs cannot be derived here.

ARC-AGI-2 (abstract reasoning, non-coding) — cost per task published

Board: arcprize.org/leaderboard. Top 10 of the verified rows by ARC-AGI-2 score ÷ cost per task (effort level shown; the ratio systematically favors low-effort configs of cheap models):

Model (effort)ARC-AGI-2Cost/taskScore% per $
DeepSeek V4 Flash 0731 (low)46.0%$0.0212190.5
DeepSeek V4 Flash 0731 (max)61.4%$0.0421461.9
DeepSeek V4 Flash 0731 (high)56.0%$0.0451244.4
Gemini 3.7 Flash (low)52.9%$0.079669.6
Gemini 3.7 Flash (medium)63.7%$0.116549.1
GPT-5.6 Luna 2026-07-30 (high)33.3%$0.063528.6
Gemini 3.7 Flash (high)84.6%$0.249339.8
GPT-5.6 Luna 2026-07-30 (max)59.6%$0.177336.7
GPT-5.6 Luna 2026-07-30 (xhigh)43.5%$0.130334.6
Inkling Small (high)33.1%$0.133248.9

Reading: DeepSeek V4 Flash is the single cheapest config on the whole ARC board — 61.4% at $0.042/task (max effort) is an order of magnitude cheaper than anything scoring comparably. GLM-5.3-flash is not on the ARC verified leaderboard yet (GLM-5 at 4.9% and GLM-5.2 at 22.8% on AGI-2 are); that is a coverage gap, not a measured result. ARC-AGI-1 costs are lower still (DeepSeek V4 Flash max: 89.0% at $0.021/task).

Boards without published per-model cost — performance only

Honest gap first: two boards in the portfolio cannot produce a performance/cost ranking from their published data.

SWE-bench Pro (public)leaderboard: resolve rate ±CI only; cost appears solely as a methodology footnote (some rows ran with a capped cost/turn limit, others uncapped). Top 10 by resolve rate, cost column not published:

ModelResolve rate
Muse Spark 1.1*61.50±3.10
GPT-5.4 (xhigh)*59.10±3.56
Muse Spark*55.00±3.60
Claude Opus 4.6 (thinking)*51.90±3.61
Gemini 3.1 Pro (thinking)*46.10±3.60
Claude Opus 4.545.89±3.60
Claude 4.5 Sonnet43.60±3.60
Gemini 3 Pro Preview43.30±3.60
Claude 4 Sonnet42.70±3.59
GPT-5 (High)41.78±3.49

* = mini-swe-agent harness; DeepSeek’s entry is V3.2 at 15.56±2.63 (rank ~19), GLM’s is GLM-4.6 at 9.67±2.15 — both older generations, so this board currently under-represents the two families this post tracks.

Vending-Bench 2 (Andon Labs)board: score is final bank balance (± across 5 runs); a Score-vs-cost-per-run chart exists, but per-row cost values are only rendered graphically, so no cost column is transcribed here. Top 10 by money balance:

ModelBalance (avg of 5 runs)
GPT-6 Astra$15,514.70 ± $1,074
Claude Opus 5$11,181.87 ± $2,094
Claude Opus 4.7$10,936.76 ± $1,181
GPT-5.6 Sol$9,619.37 ± $1,338
Grok 4.6$9,047.03 ± $1,604
GLM-5.2$8,313.78 ± $1,084
GLM-5.3$8,163.61 ± $787
Claude Opus 4.6$8,017.59 ± $1,367
GPT-5.5$7,523.84 ± $1,346
GPT-5.6 Terra$7,343.21 ± $373

DeepSeek and GLM-5.3-flash rows sit below the fold (53 more rows) and could not be verified from the static page; their absence above is not evidence of absence.

Artificial Analysis indexes — cost-per-task defined, bulk table not fetchable

AA’s Intelligence Index v4.3.2 (10 evals including Terminal-Bench 4.0, HLE, GDPval-AA, SciCode) publishes a weighted “Cost per Intelligence Index task” and an Index-vs-cost Pareto chart, and the Coding Agent Index (DeepSWE v1.1 + TB 4.0 + SWE-Atlas-QnA) publishes cost-per-task Pareto charts. The per-model values live on individual model pages that render charts client-side; bulk transcription was not reproducible from static fetches, so no top-10 is produced here. One verified example row: GLM-5.3-flash — Intelligence Index 42, $0.25 per index task ($0.15/M input, $0.50/M output, 83% cache discount; $280.28 total to evaluate), 320B total/18B active params, 1M context, MIT license. Per-model pulls are scripted in the rerun template below.

Coverage of the two tracked families

BoardGLM-5.3-flashDeepSeek
DeepSWE v1.1✅ 63% @ $0.24✅ v4-flash 53% @ $0.46 · v4-pro 63% @ $1.67
ARC-AGI-2❌ not yet run (GLM-5/5.2 are)✅ V4 Pro + V4 Flash, full cost rows
Terminal-Bench 4.0⚠️ not on board (GLM-5.3 is); run by AA inside its index❌ not on 14-row board
Vending-Bench 2⚠️ GLM-5.2/5.3 verified; flash unverified⚠️ unverified (below fold)
SWE-bench Pro❌ (GLM-4.6)✅ V3.2 (older gen)
AA Indexes✅ model page + index✅ tracked

If only three boards are tracked, DeepSWE + the two AA indexes carry both families across coding and mixed domains; ARC-AGI-2 is the non-coding board with the best cost granularity for DeepSeek.

Reading rules (the caveats that bite)

  • Ratio ≠ preference. Score-per-dollar is reproducible but unit-dependent; its level sets are hyperbolas through the origin and it silently rewards low scorers (Luna at 17.3% topping TB 4.0 is the illustration). Use the ratio to find candidates, the Pareto set to make claims.
  • Cost columns are dated receipts. Two of ten Terminal-Bench 3.0 rows were billed at price cards later withdrawn (TokenCost’s audit), repricing one run from $30.19 to $6.04 per solved task. Compare rows only after checking each run’s date against the vendor’s pricing history.
  • Agent workloads are cache bills. On TB 3.0, cached input was 92–98% of tokens on every run and 61–73% of the bill on the reproducible rows. Cost charts computed without caching (Vending-Bench 2’s) overstate relative costs for models with cheap cache reads — my interpretation, flagged as such.
  • Harness varies on some boards. TB 4.0 rows use five different agents; DeepSWE holds mini-swe-agent fixed. Cross-board comparisons inherit both.
  • Effort levels dominate ratio rankings. ARC’s top-10 is all low-effort-flash configs; effort-normalized comparisons need the Pareto view.

Rerun procedure

New-model reruns follow the operational template templates/research/cost-frontier-benchmarks-rerun.md in this repo: fetch each board, extract (score, cost) rows, recompute the ratio top-10 and the Pareto frontier, check the new model’s coverage, and emit a dated snapshot. That template is the operational companion to this post.

Methodology & Sources

All figures fetched live on 2026-09-17 (DeepSWE leaderboard self-reported as updated 2026-09-03). Ratio and Pareto computations are derivations from those published (score, cost) pairs; arithmetic shown in the tables is reproducible by sorting and a dominance sweep. Boards’ own numbers are vendor-published results — Datacurve, ARC Prize, Laude Institute, Scale AI, Andon Labs, and Artificial Analysis each publish their own runs.

Companions