Graff evaluations.
Recorded coding-agent results, from terminal tasks to verified repository patches. Inspect the outcomes and the conditions behind them.
Original report ·
Read the experiment01 / TASK COMPLETION
21 terminal tasks.
Selected Terminal-Bench 2.1 tasks.
The final environment is graded with public tests.
These Graff runs used additional evaluation instructions and a local Docker runtime.
Graff / Grok 4.6
20 / 2195.2%
$0.31 / completed task
$6.2335 recorded spend · failures included
Graff / Kimi K3
17 / 2181.0%
$0.21 / completed task
$3.6463 recorded spend · failures included
02 / PATCH VERIFICATION
One pass. Eight misses.
Nine DeepSWE tasks.
Submitted patches face the project’s real verifier.
Both Graff runs passed the KaTeX multicolumn-array-spans task. The other eight did not pass. Patch-application failures and test failures remain part of the result.
Graff / Grok 4.6
1 / 911.1%
$25.56 / completed task
$25.5616 recorded spend · failures included
Graff / Kimi K3
1 / 911.1%
$21.91 / completed task
$21.9148 recorded spend · failures included
03 / THE WIDER FIELD
Compare the configurations.
The published FrontierHarness field contains twelve Kimi K3 configurations. Graff used different instructions and runtime conditions; those differences belong alongside the numbers.
Costs include failed attempts. ≥ marks a lower bound where some task costs are missing.
Recorded Graff configurations
Extra terminal instructions · local Docker runtime · historical list-price estimates
| Configuration | Completed | Recorded USD | USD / pass | Costed tasks | Median attempt |
|---|---|---|---|---|---|
| Graff / Grok 4.6 | 20/21 | $6.2335 | $0.31 | 21/21 | Not reported |
| Graff / Kimi K3 | 17/21 | $3.6463 | $0.21 | 21/21 | Not reported |
Published Kimi K3 field
Shared published runtime and instructions · first-turn-cold pricing
| Configuration | Completed | Recorded USD | USD / pass | Costed tasks | Median attempt |
|---|---|---|---|---|---|
| Codex | 15/21 | $4.8584 | $0.32 | 21/21 | 6m 51s |
| DSH Creator | 15/21 | ≥ $3.3749 | ≥ $0.22 | 18/21 | 7m 33s |
| Claude Code | 16/21 | $15.5788 | $0.97 | 21/21 | 9m 38s |
| Pi | 16/21 | $5.4599 | $0.34 | 21/21 | 8m 06s |
| DSH PTC | 16/21 | ≥ $3.0198 | ≥ $0.19 | 18/21 | 7m 55s |
| DSH Standard | 16/21 | ≥ $8.1457 | ≥ $0.51 | 19/21 | 7m 04s |
| Oh My Pi | 16/21 | $5.4428 | $0.34 | 21/21 | 7m 24s |
| Kimi Code | 14/21 | $9.3918 | $0.67 | 21/21 | 7m 56s |
| DSH Minimal | 15/21 | ≥ $6.2995 | ≥ $0.42 | 17/21 | 6m 55s |
| Exo Harness | 16/21 | $6.4250 | $0.40 | 21/21 | 6m 51s |
| OpenCode | 15/21 | $3.4771 | $0.23 | 21/21 | 7m 43s |
| Hermes | 13/21 | $11.3272 | $0.87 | 21/21 | 9m 08s |
Costs include failed attempts. ≥ marks a lower bound where some task costs are missing.
Recorded Graff configurations
Extra terminal instructions · local Docker runtime · historical list-price estimates
| Configuration | Completed | Recorded USD | USD / pass | Costed tasks | Median attempt |
|---|---|---|---|---|---|
| Graff / Grok 4.6 | 1/9 | $25.5616 | $25.56 | 9/9 | Not reported |
| Graff / Kimi K3 | 1/9 | $21.9148 | $21.91 | 9/9 | Not reported |
Published Kimi K3 field
Shared published runtime and instructions · first-turn-cold pricing
| Configuration | Completed | Recorded USD | USD / pass | Costed tasks | Median attempt |
|---|---|---|---|---|---|
| Codex | 5/9 | $64.5075 | $12.90 | 9/9 | 43m 07s |
| DSH Creator | 4/9 | $59.0383 | $14.76 | 9/9 | 37m 12s |
| Claude Code | 3/9 | ≥ $332.8210 | ≥ $110.94 | 8/9 | 59m 56s |
| Pi | 2/9 | $38.3334 | $19.17 | 9/9 | 39m 11s |
| DSH PTC | 2/9 | $79.3743 | $39.69 | 9/9 | 50m 21s |
| DSH Standard | 2/9 | ≥ $54.1143 | ≥ $27.06 | 8/9 | 47m 16s |
| Oh My Pi | 1/9 | $75.2314 | $75.23 | 9/9 | 49m 57s |
| Kimi Code | 3/9 | $52.6087 | $17.54 | 9/9 | 44m 29s |
| DSH Minimal | 2/9 | $73.9058 | $36.95 | 9/9 | 41m 25s |
| Exo Harness | 0/9 | $10.2977 | No passes | 9/9 | 15m 55s |
| OpenCode | 0/9 | ≥ $45.1881 | No passes | 8/9 | 49m 05s |
| Hermes | 2/9 | ≥ $32.2374 | ≥ $16.12 | 7/9 | 28m 48s |
Combined totals add the terminal and DeepSWE slices. Costs include failed attempts. ≥ marks a lower bound where some task costs are missing.
Recorded Graff configurations
Extra terminal instructions · local Docker runtime · historical list-price estimates
| Configuration | Completed | Recorded USD | USD / pass | Costed tasks | Median attempt |
|---|---|---|---|---|---|
| Graff / Grok 4.6 | 21/30 | $31.7951 | $1.51 | 30/30 | Not reported |
| Graff / Kimi K3 | 18/30 | $25.5611 | $1.42 | 30/30 | Not reported |
Published Kimi K3 field
Shared published runtime and instructions · first-turn-cold pricing
| Configuration | Completed | Recorded USD | USD / pass | Costed tasks | Median attempt |
|---|---|---|---|---|---|
| Codex | 20/30 | $69.3659 | $3.47 | 30/30 | 9m 22s |
| DSH Creator | 19/30 | ≥ $62.4132 | ≥ $3.28 | 27/30 | 12m 29s |
| Claude Code | 19/30 | ≥ $348.3998 | ≥ $18.34 | 29/30 | 18m 29s |
| Pi | 18/30 | $43.7934 | $2.43 | 30/30 | 10m 39s |
| DSH PTC | 18/30 | ≥ $82.3941 | ≥ $4.58 | 27/30 | 8m 54s |
| DSH Standard | 18/30 | ≥ $62.2600 | ≥ $3.46 | 27/30 | 9m 53s |
| Oh My Pi | 17/30 | $80.6742 | $4.75 | 30/30 | 9m 56s |
| Kimi Code | 17/30 | $62.0005 | $3.65 | 30/30 | 9m 20s |
| DSH Minimal | 17/30 | ≥ $80.2052 | ≥ $4.72 | 26/30 | 14m 54s |
| Exo Harness | 16/30 | $16.7227 | $1.05 | 30/30 | 10m 04s |
| OpenCode | 15/30 | ≥ $48.6652 | ≥ $3.24 | 29/30 | 11m 41s |
| Hermes | 15/30 | ≥ $43.5646 | ≥ $2.90 | 28/30 | 13m 06s |
04 / HOW TO READ THE RESULTS
Keep the conditions attached.
Checking your own patch? Follow the verification workflow to retain commands, outcomes and a reviewable diff.
Results are repository-reported evaluations checked against the tracked summaries and protocol at revision b4bd80c, not a fresh rerun or an independent audit. The Graff terminal runs shown here used additional evaluation instructions. Models and execution environments differ across some rows. Published Exo costs are taken from FrontierHarness revision 8f11b13; the terminal subtotal is calculated from its public task-level records.
What does each task set measure?
The selected 21 Terminal-Bench 2.1 tasks grade the environment left behind by the agent. The nine DeepSWE tasks grade submitted patches against real project verifiers. These results cover those slices, not the full Terminal-Bench leaderboard.
Combined totals add both slices: Graff / Grok completes 21/30 and Graff / Kimi completes 18/30. Keep the individual slices visible when interpreting those totals.
Why are Graff and the published field separate?
Graff’s later terminal runs received BENCH_APPEND: additional evaluation guidance about final files and persistent services. They used a local Docker runtime. The published Kimi K3 field used a shared setup with prepared VM checkpoints and did not receive Graff’s append.
Grok also differs from the published field’s model. The recorded outcomes describe those complete configurations; they do not isolate the effect of switching harnesses alone.
How are cost and time calculated?
Cost per pass is recorded spend divided by completed tasks, including spending on failures. Graff uses historical evaluator list-price estimates. The published field uses cost_first_cold_usd, which reprices first-turn cache reads. These are different accounting methods, not current provider price quotes.
Where task costs are missing, ≥ denotes a lower bound. No passes means the ratio is undefined, not free. Median attempt time includes successful and failed attempts where timings are available. Graff medians are not reported; summary maxima are not medians.
Can I inspect the underlying evidence?
These are repository-reported results, not a new rerun or independent audit. Graff records are pinned to b4bd80c; the published board is pinned to 8f11b13. Publication and page-update dates do not imply new evaluation runs.
- FrontierHarness results and suite definitions
- Evaluation protocol and comparison limitations
- Graff / Grok terminal result summary
- Graff / Kimi K3 terminal result summary
- Published FrontierHarness task costs and Exo results
- Published cost accounting and comparison method
- Graff / Grok DeepSWE cost summary
- Graff / Kimi K3 DeepSWE cost summary
THE FULL WRITE-UP
What happened inside the runs?
Task failures, model differences, pricing methods, and the reasoning behind the comparisons.
Read the experiment and methodology