bang for buck, not bang for ego

Coding models ranked by dollars per solved task.

Measured cost per solved task in the waterfall.sh harness. Not list price. Not vendor SWE-bench.

How to read this

Value is quality per dollar spent on a solved task, normalized so the board leader is 100.

The quality ceiling is useful, but it is not the same thing as the best default.

Snapshot costs are public priors; harness costs include retries and failed attempts as measured runs replace them.

Cache reads can change Fable economics, so Waterfall will not recommend Fable on cost unless the harness actually cached the stable prefix.

Live table

Value sort, snapshot dated 2026-09-03.

Rank Model List in / out Cache read Cost / attempt Best for Updated
1DeepSeek V4 Flashmedium · snapshot72$0.22 / $0.66none$0.19pending62%100.0High-volume first pass2026-09-03
2MiniMax M3medium · snapshot74$0.30 / $1.20none$0.35pending70%55.8Cheap capable implement2026-09-03
3GLM-5.3max · snapshot86$1.40 / $4.40none$0.68pending78%33.4Frontier-cluster value2026-09-03
4Grok 4.6medium · snapshot88$2.00 / $6.00none$0.84pending80%27.6Default implementer2026-09-03
5Kimi K3max · snapshot87$3.00 / $15.00none$0.94pending82%24.4Frontend / UI agents2026-09-03
6Claude Opus 5high · snapshot93$5.00 / $25.00none$2.34pending90%10.5Quality floor / review2026-09-03
7Claude Fable 5.1high · snapshot96$10.00 / $50.00$0.25$2.72pending93%9.3Escalate after Opus fails2026-09-03
8Claude Fable 5.1max · snapshot98$10.00 / $50.00$0.25$3.76pending95%6.9Ceiling only2026-09-03
9GPT-5.6 Terramax · snapshot89$2.00 / $12.00none$4.95pending84%4.7OpenAI everyday agent2026-09-03
10GPT-5.6 Solmax · snapshot92$4.00 / $20.00none$8.39pending89%2.9Tool-heavy agents2026-09-03

Coverage: borrowed scores, not measured

Loading catalog coverage.

Everything above has a measured or dated cost per solved task. Everything below does not. These rows borrow Artificial Analysis coding_index (or intelligence_index, or rescaled Design Arena Elo) from OpenRouter's catalog so the other two hundred models are at least visible. No value score, because the formula needs a measured cost per solved task. Prices are list, per 1M tokens.

# Model Borrowed quality Source List in / out Context Efforts
Coverage needs the JSON feed; see /api/leaderboard.json.

Waterfall default stack

Cheap first. Independent review. Expensive only after a blocking failure.

01 · draftDeepSeek Flashlow effort
02 · implementGrok / GLMmedium effort
03 · hardenOpus 5high effort
04 · escalateFable 5.1only after failure

Methodology

Score

value_raw = quality / max(cost_per_solved, 0.01)

value = 100 × value_raw / max(value_raw on board)

Harness quality is the equal-weight mean of pass rates across task categories. A solve means the declared tests or blocking reviewer gate passed on the first submitted attempt.

Source and sample

snapshot-2026-09-03 means a prior compiled from public September 2026 reporting and list prices. It is not a Waterfall measurement, so cost per attempt and n remain pending.

harness means committed task-level records from data/runs/*.jsonl, including model, effort, tokens, cache reads, cost, wall time, pass/fail, promotions, and reviewer verdict.

Context, not borrowed results

SWE-bench, Artificial Analysis, and DeepSWE are useful external context. Their scores are never relabeled as Waterfall harness numbers.

The shipped suite is intentionally small: 5 one-file bugfixes, 5 multi-file features, 3 review-only diffs, and 2 long-horizon finish-the-PR tasks.

coverage rows are the exception that proves the rule: borrowed AA and Design Arena scores pulled from OpenRouter's catalog for models the harness has not run, kept in their own table with their own field names so they cannot be mistaken for measurements.

Snapshot date: 2026-09-03. Prices move; this row is dated. Rebuild the feeds with waterfall leaderboard --publish.