Benchmark Cost Frontiers — Top-10 Performance-per-Dollar Across Coding and Non-Coding Boards (2026-09-20)
A dated snapshot ranking model configurations by performance per dollar on five boards (DeepSWE v1.1, Terminal-Bench 4.0, ARC-AGI-2, SWE-bench Pro, Vending-Bench 2), with published costs listed per row, the DeepSWE Pareto frontier computed two ways, and a coverage check for GLM-5.3-Flash and DeepSeek. Rerun procedure lives in templates/research/cost-frontier-benchmarks-rerun.md.
Snapshot: data fetched live on 2026-09-17 (DeepSWE board last updated by its
operator on 2026-09-03; other boards as fetched the same day). Costs are the
boards’ own published figures on the day each run happened — read them as dated
receipts, not current quotes. Rerun checklist:
templates/research/cost-frontier-benchmarks-rerun.md in this repo.
What “top 10 by performance/cost” means here
Three readings exist, and they disagree. This snapshot separates them instead of blending:
- Ratio ranking —
score ÷ published cost per task. What the tables below show. Simple and reproducible, but the ratio has known flaws (see caveats): it rewards low scorers on cheap boards, and its level sets are hyperbolas, not a preference. - Pareto frontier — the non-dominated set: a config stays only if no cheaper-or-equal config has an equal-or-better score. This is the preference-free answer to “which models should exist on the menu.”
- Cost per solved task —
cost ÷ (score × tasks). A different lens that reorders boards; it flatters models that attempt cheaply and fail cheaply.
All costs are the benchmark operators’ published figures (vendor-run for their own benchmarks; independent operators for cross-model boards). Where a board does not publish per-model cost, the table says so rather than inventing one.
DeepSWE v1.1 (agentic coding, 113 tasks) — cost per task published
Board: deepswe.datacurve.ai, 21 “best” configurations, updated 2026-09-03. Metric: Pass@1 ± CI under a fixed mini-swe-agent harness. Top 10 of 21 by Pass@1 per dollar:
| Model (effort) | Pass@1 | Avg cost/task | Pass@1 per $ |
|---|---|---|---|
| GLM-5.3-flash (max) | 63%±4% | $0.24 | 262.5 |
| DeepSeek V4 Flash (max) | 53%±4% | $0.46 | 115.2 |
| GPT-5.6 Luna (max) | 67%±4% | $0.61 | 109.8 |
| DeepSeek V4 Pro (max) | 63%±6% | $1.67 | 37.7 |
| Gemini 3.7 Flash (medium) | 65%±3% | $2.03 | 32.0 |
| Gemini 3.8 Flash (high) | 74%±1% | $2.36 | 31.4 |
| Gemini 3.6 Flash (high) | 47%±4% | $2.21 | 21.3 |
| Grok 4.6 (medium) | 67%±2% | $3.45 | 19.4 |
| GLM-5.3 (max) | 69%±3% | $3.99 | 17.3 |
| Qwen3.8-max (xhigh) | 57%±3% | $3.73 | 15.3 |
The same board under the strict Pareto reading keeps only three configs:
| Frontier | Pass@1 | Cost/task |
|---|---|---|
| GLM-5.3-flash (max) | 63%±4% | $0.24 |
| GPT-5.6 Luna (max) | 67%±4% | $0.61 |
| Gemini 3.8 Flash (high) | 74%±1% | $2.36 |
Everything else is dominated — including the expensive tier: Claude Opus 5 (74% at $11.84) and GPT-6 Astra (74% at $6.52) are deleted by Gemini 3.8 Flash (74% at $2.36, equal score, lower cost); Claude Fable 5 (70% at $13.41) by the same point. Arithmetic (cost ascending, keep strictly better running-max score) was cross-checked against an independent O(n²) dominance test — identical set. Within the ±4-point confidence intervals, GLM-5.3-flash vs GPT-5.6 Luna overlap, so “dominated” here means point estimates, not proven gaps.
Terminal-Bench 4.0 (agentic terminal, 66 tasks) — total run cost published
Board: tbench.ai. Publishes resolution rate (±95% CI) and total COST and TOKENS per run; the ratio below divides by the published whole-board run cost, so it is a derived ranking. Rows use different agent harnesses (Codex, Claude Code, Grok Build, mini-SWE-agent), which the board itself flags per row — treat harness differences as part of the measurement. 14 rows exist; top 10:
| Model (effort) | Agent | Resolution | Run cost | Res.% per $1k |
|---|---|---|---|---|
| GPT-5.6 Luna (max) | Codex | 17.3%±2.8% | $0.3k | 57.7 |
| GPT-6 Astra (max) | Codex | 58.2%±2.8% | $3.3k | 17.6 |
| GLM-5.3 (max) | Claude Code | 41.8%±3.2% | $2.7k | 15.5 |
| GPT-5.6 Sol (max) | Codex | 37.3%±3.8% | $2.5k | 14.9 |
| GPT-5.6 Terra (max) | Codex | 21.5%±3.3% | $1.7k | 12.6 |
| Gemini 3.8 Flash (high) | mini-SWE-agent | 19.1%±3.4% | $1.8k | 10.6 |
| Fable 5.1 (max) | Claude Code | 57.9%±3.8% | $6.2k | 9.3 |
| Opus 5 (xhigh) | Claude Code | 53.9%±3.2% | $6.1k | 8.8 |
| Gemini 3.7 Flash (high) | mini-SWE-agent | 11.2%±2.4% | $1.3k | 8.6 |
| Fable 5 (max) | Claude Code | 44.5%±3.8% | $7.3k | 6.1 |
Notes: no GLM-5.3-flash and no DeepSeek row on the current board (GLM-5.3 is the family’s representative). GPT-5.6 Luna’s ratio is inflated by a very low score — ratio rewards doing a cheap subset (the same trap TokenCost documented for cost-per-solved-task on TB 3.0). Attempts-per-task for TB 4.0 runs are not published on the board I could fetch, so absolute per-task costs cannot be derived here.
ARC-AGI-2 (abstract reasoning, non-coding) — cost per task published
Board: arcprize.org/leaderboard. Top 10 of the verified rows by ARC-AGI-2 score ÷ cost per task (effort level shown; the ratio systematically favors low-effort configs of cheap models):
| Model (effort) | ARC-AGI-2 | Cost/task | Score% per $ |
|---|---|---|---|
| DeepSeek V4 Flash 0731 (low) | 46.0% | $0.021 | 2190.5 |
| DeepSeek V4 Flash 0731 (max) | 61.4% | $0.042 | 1461.9 |
| DeepSeek V4 Flash 0731 (high) | 56.0% | $0.045 | 1244.4 |
| Gemini 3.7 Flash (low) | 52.9% | $0.079 | 669.6 |
| Gemini 3.7 Flash (medium) | 63.7% | $0.116 | 549.1 |
| GPT-5.6 Luna 2026-07-30 (high) | 33.3% | $0.063 | 528.6 |
| Gemini 3.7 Flash (high) | 84.6% | $0.249 | 339.8 |
| GPT-5.6 Luna 2026-07-30 (max) | 59.6% | $0.177 | 336.7 |
| GPT-5.6 Luna 2026-07-30 (xhigh) | 43.5% | $0.130 | 334.6 |
| Inkling Small (high) | 33.1% | $0.133 | 248.9 |
Reading: DeepSeek V4 Flash is the single cheapest config on the whole ARC board — 61.4% at $0.042/task (max effort) is an order of magnitude cheaper than anything scoring comparably. GLM-5.3-flash is not on the ARC verified leaderboard yet (GLM-5 at 4.9% and GLM-5.2 at 22.8% on AGI-2 are); that is a coverage gap, not a measured result. ARC-AGI-1 costs are lower still (DeepSeek V4 Flash max: 89.0% at $0.021/task).
Boards without published per-model cost — performance only
Honest gap first: two boards in the portfolio cannot produce a performance/cost ranking from their published data.
SWE-bench Pro (public) — leaderboard: resolve rate ±CI only; cost appears solely as a methodology footnote (some rows ran with a capped cost/turn limit, others uncapped). Top 10 by resolve rate, cost column not published:
| Model | Resolve rate |
|---|---|
| Muse Spark 1.1* | 61.50±3.10 |
| GPT-5.4 (xhigh)* | 59.10±3.56 |
| Muse Spark* | 55.00±3.60 |
| Claude Opus 4.6 (thinking)* | 51.90±3.61 |
| Gemini 3.1 Pro (thinking)* | 46.10±3.60 |
| Claude Opus 4.5 | 45.89±3.60 |
| Claude 4.5 Sonnet | 43.60±3.60 |
| Gemini 3 Pro Preview | 43.30±3.60 |
| Claude 4 Sonnet | 42.70±3.59 |
| GPT-5 (High) | 41.78±3.49 |
* = mini-swe-agent harness; DeepSeek’s entry is V3.2 at 15.56±2.63 (rank ~19), GLM’s is GLM-4.6 at 9.67±2.15 — both older generations, so this board currently under-represents the two families this post tracks.
Vending-Bench 2 (Andon Labs) — board: score is final bank balance (± across 5 runs); a Score-vs-cost-per-run chart exists, but per-row cost values are only rendered graphically, so no cost column is transcribed here. Top 10 by money balance:
| Model | Balance (avg of 5 runs) |
|---|---|
| GPT-6 Astra | $15,514.70 ± $1,074 |
| Claude Opus 5 | $11,181.87 ± $2,094 |
| Claude Opus 4.7 | $10,936.76 ± $1,181 |
| GPT-5.6 Sol | $9,619.37 ± $1,338 |
| Grok 4.6 | $9,047.03 ± $1,604 |
| GLM-5.2 | $8,313.78 ± $1,084 |
| GLM-5.3 | $8,163.61 ± $787 |
| Claude Opus 4.6 | $8,017.59 ± $1,367 |
| GPT-5.5 | $7,523.84 ± $1,346 |
| GPT-5.6 Terra | $7,343.21 ± $373 |
DeepSeek and GLM-5.3-flash rows sit below the fold (53 more rows) and could not be verified from the static page; their absence above is not evidence of absence.
Artificial Analysis indexes — cost-per-task defined, bulk table not fetchable
AA’s Intelligence Index v4.3.2 (10 evals including Terminal-Bench 4.0, HLE, GDPval-AA, SciCode) publishes a weighted “Cost per Intelligence Index task” and an Index-vs-cost Pareto chart, and the Coding Agent Index (DeepSWE v1.1 + TB 4.0 + SWE-Atlas-QnA) publishes cost-per-task Pareto charts. The per-model values live on individual model pages that render charts client-side; bulk transcription was not reproducible from static fetches, so no top-10 is produced here. One verified example row: GLM-5.3-flash — Intelligence Index 42, $0.25 per index task ($0.15/M input, $0.50/M output, 83% cache discount; $280.28 total to evaluate), 320B total/18B active params, 1M context, MIT license. Per-model pulls are scripted in the rerun template below.
Coverage of the two tracked families
| Board | GLM-5.3-flash | DeepSeek |
|---|---|---|
| DeepSWE v1.1 | ✅ 63% @ $0.24 | ✅ v4-flash 53% @ $0.46 · v4-pro 63% @ $1.67 |
| ARC-AGI-2 | ❌ not yet run (GLM-5/5.2 are) | ✅ V4 Pro + V4 Flash, full cost rows |
| Terminal-Bench 4.0 | ⚠️ not on board (GLM-5.3 is); run by AA inside its index | ❌ not on 14-row board |
| Vending-Bench 2 | ⚠️ GLM-5.2/5.3 verified; flash unverified | ⚠️ unverified (below fold) |
| SWE-bench Pro | ❌ (GLM-4.6) | ✅ V3.2 (older gen) |
| AA Indexes | ✅ model page + index | ✅ tracked |
If only three boards are tracked, DeepSWE + the two AA indexes carry both families across coding and mixed domains; ARC-AGI-2 is the non-coding board with the best cost granularity for DeepSeek.
Reading rules (the caveats that bite)
- Ratio ≠ preference. Score-per-dollar is reproducible but unit-dependent; its level sets are hyperbolas through the origin and it silently rewards low scorers (Luna at 17.3% topping TB 4.0 is the illustration). Use the ratio to find candidates, the Pareto set to make claims.
- Cost columns are dated receipts. Two of ten Terminal-Bench 3.0 rows were billed at price cards later withdrawn (TokenCost’s audit), repricing one run from $30.19 to $6.04 per solved task. Compare rows only after checking each run’s date against the vendor’s pricing history.
- Agent workloads are cache bills. On TB 3.0, cached input was 92–98% of tokens on every run and 61–73% of the bill on the reproducible rows. Cost charts computed without caching (Vending-Bench 2’s) overstate relative costs for models with cheap cache reads — my interpretation, flagged as such.
- Harness varies on some boards. TB 4.0 rows use five different agents; DeepSWE holds mini-swe-agent fixed. Cross-board comparisons inherit both.
- Effort levels dominate ratio rankings. ARC’s top-10 is all low-effort-flash configs; effort-normalized comparisons need the Pareto view.
Rerun procedure
New-model reruns follow the operational template
templates/research/cost-frontier-benchmarks-rerun.md in this repo: fetch each
board, extract (score, cost) rows, recompute the ratio top-10 and the Pareto
frontier, check the new model’s coverage, and emit a dated snapshot.
That template is the operational companion to this post.
Methodology & Sources
All figures fetched live on 2026-09-17 (DeepSWE leaderboard self-reported as updated 2026-09-03). Ratio and Pareto computations are derivations from those published (score, cost) pairs; arithmetic shown in the tables is reproducible by sorting and a dominance sweep. Boards’ own numbers are vendor-published results — Datacurve, ARC Prize, Laude Institute, Scale AI, Andon Labs, and Artificial Analysis each publish their own runs.
- DeepSWE v1.1 leaderboard — https://deepswe.datacurve.ai/ (v1 methodology: https://deepswe.datacurve.ai/blog/deepswe)
- Terminal-Bench 4.0 leaderboard — https://www.tbench.ai/leaderboard
- ARC-AGI leaderboard (AGI-1/2/3, cost-per-task rows) — https://arcprize.org/leaderboard
- SWE-bench Pro public leaderboard — https://labs.scale.com/leaderboard/swe_bench_pro_public
- Vending-Bench 2 — https://andonlabs.com/evals/vending-bench-2
- Artificial Analysis Intelligence Index (per-model cost/task) — https://artificialanalysis.ai/models/glm-5-3-flash
- Artificial Analysis Coding Agent Index — https://artificialanalysis.ai/agents/coding-agents
- TokenCost, “Terminal-Bench 3.0 Cost: $6 to $200 a Solved Task” (2026-08-17) — https://tokencost.app/blog/terminal-bench-3-cost-per-solved-task