cost & model
Where tokens and dollars go across every model and suite in the harness. Cost is the primary efficiency signal; tokens are the budget. Budget alerts flag any model running over the per-task or per-run thresholds.
$70.2
(6 models have no cost data)
57.86M
325
11
$0.40
| Model | Tasks | Pass rate | Total cost | Total tokens | Cost / solved | Avg time | Top tools |
|---|---|---|---|---|---|---|---|
| moonshot-v1-128kkimi | 43 | 7% | $63.2 | 31.59M | $21.1 | 25m 1s | n/a |
| agy+MiniMax-M3+deepseek-v4-proconsensus | 36 | 100% | $2.6 | 7.88M | $0.07 | 3m 4s | n/a |
| deepseek-v4-prodeepseek | 59 | 39% | $2.0 | 6.36M | $0.09 | 2m 45s | n/a |
| MiniMax-M3minimax | 48 | 31% | $1.7 | 5.00M | $0.11 | 2m 34s | n/a |
| qwen3.8-max-previewqwen | 20 | 70% | $0.72 | 2.27M | $0.05 | 2m 55s | n/a |
| deepseek-v4-pro+qwen3.8-max-previewconsensus | 10 | 100% | n/a | 2.35M | n/a | 3m 29s | n/a |
| defaultinternal | 49 | 82% | n/a | 2.17M | n/a | 1m 37s | n/a |
| kimi-k3kimi | 2 | 100% | n/a | 189.8k | n/a | 2m 47s | n/a |
| nvidia/nemotron-3-ultra-550b-a55b:freeopenrouter | 1 | 100% | n/a | 45.1k | n/a | 48.9s | n/a |
| agyexternal | 47 | 64% | n/a | n/a | n/a | 2m 24s | n/a |
| agyconsensus | 10 | 0% | n/a | n/a | n/a | 4.2s | n/a |
| Suite | Models | Runs | Passed | Total cost | Total tokens | Total time |
|---|---|---|---|---|---|---|
| deep-swe | 1 | 43 | 3/43 | $63.2 | 31.59M | 17h 55m |
| harder | 4 | 145 | 73/145 | $5.2 | 15.82M | 7h 21m |
| hard-targeting | 7 | 76 | 66/76 | $1.8 | 8.98M | 3h 3m |
| fast | 4 | 22 | 2/22 | $0.01 | 88.5k | 2m 59s |
| harder-v2 | 1 | 33 | 30/33 | n/a | 1.39M | 36m 27s |
| swebench-lite | 1 | 6 | 0/6 | n/a | n/a | 24m 3s |
| Complexity | Runs | Passed | Total cost | Total tokens | Avg time |
|---|---|---|---|---|---|
| complex | 101 | 54/101 | $62.0 | 37.62M | 9m 53s |
| medium | 190 | 112/190 | $7.6 | 18.36M | 3m 54s |
| simple | 34 | 8/34 | $0.60 | 1.88M | 45.8s |
Methodology
Cost and token data comes from the model provider responses (where the API exposes usage). External CLIs like agy (PTY-based, no metrics parser) and providers that don't return cost in the chat response (e.g. z.ai GLM) show as n/a. Pass-rate is computed from the harness server's own evaluation, not the model self-report.
Budget thresholds: per-task cost > $0.50 triggers a warn, > $2.00 an alert. Total run cost > $5.00 also flags as alert.