Graff evaluations.

Recorded coding-agent results, from terminal tasks to verified repository patches. Inspect the outcomes and the conditions behind them.

Original report ·

Read the experiment

01 / TASK COMPLETION

21 terminal tasks.

Selected Terminal-Bench 2.1 tasks.
The final environment is graded with public tests.

These Graff runs used additional evaluation instructions and a local Docker runtime.

Completed tasks No recorded pass

Graff / Grok 4.6

20 / 2195.2%

$0.31 / completed task

$6.2335 recorded spend · failures included

Graff / Kimi K3

17 / 2181.0%

$0.21 / completed task

$3.6463 recorded spend · failures included

Each segment represents one task in the total, not a specific task identity. Costs use the evaluator’s historical list-price estimates.

02 / PATCH VERIFICATION

One pass. Eight misses.

Nine DeepSWE tasks.
Submitted patches face the project’s real verifier.

Both Graff runs passed the KaTeX multicolumn-array-spans task. The other eight did not pass. Patch-application failures and test failures remain part of the result.

Graff / Grok 4.6

1 / 911.1%

$25.56 / completed task

$25.5616 recorded spend · failures included

Graff / Kimi K3

1 / 911.1%

$21.91 / completed task

$21.9148 recorded spend · failures included

A near-pass is still a failure. A missing verifier result never becomes an inferred success.

03 / THE WIDER FIELD

Compare the configurations.

The published FrontierHarness field contains twelve Kimi K3 configurations. Graff used different instructions and runtime conditions; those differences belong alongside the numbers.

Costs include failed attempts. ≥ marks a lower bound where some task costs are missing.

Recorded Graff configurations

Extra terminal instructions · local Docker runtime · historical list-price estimates

21 terminal tasks — Graff
ConfigurationCompletedRecorded USDUSD / passCosted tasksMedian attempt
Graff / Grok 4.620/21$6.2335$0.3121/21Not reported
Graff / Kimi K317/21$3.6463$0.2121/21Not reported

Published Kimi K3 field

Shared published runtime and instructions · first-turn-cold pricing

21 terminal tasks — published Kimi K3 configurations
ConfigurationCompletedRecorded USDUSD / passCosted tasksMedian attempt
Codex15/21$4.8584$0.3221/216m 51s
DSH Creator15/21≥ $3.3749≥ $0.2218/217m 33s
Claude Code16/21$15.5788$0.9721/219m 38s
Pi16/21$5.4599$0.3421/218m 06s
DSH PTC16/21≥ $3.0198≥ $0.1918/217m 55s
DSH Standard16/21≥ $8.1457≥ $0.5119/217m 04s
Oh My Pi16/21$5.4428$0.3421/217m 24s
Kimi Code14/21$9.3918$0.6721/217m 56s
DSH Minimal15/21≥ $6.2995≥ $0.4217/216m 55s
Exo Harness16/21$6.4250$0.4021/216m 51s
OpenCode15/21$3.4771$0.2321/217m 43s
Hermes13/21$11.3272$0.8721/219m 08s
Read the comparison analysis →

04 / HOW TO READ THE RESULTS

Keep the conditions attached.

Checking your own patch? Follow the verification workflow to retain commands, outcomes and a reviewable diff.

Results are repository-reported evaluations checked against the tracked summaries and protocol at revision b4bd80c, not a fresh rerun or an independent audit. The Graff terminal runs shown here used additional evaluation instructions. Models and execution environments differ across some rows. Published Exo costs are taken from FrontierHarness revision 8f11b13; the terminal subtotal is calculated from its public task-level records.

What does each task set measure?

The selected 21 Terminal-Bench 2.1 tasks grade the environment left behind by the agent. The nine DeepSWE tasks grade submitted patches against real project verifiers. These results cover those slices, not the full Terminal-Bench leaderboard.

Combined totals add both slices: Graff / Grok completes 21/30 and Graff / Kimi completes 18/30. Keep the individual slices visible when interpreting those totals.

Why are Graff and the published field separate?

Graff’s later terminal runs received BENCH_APPEND: additional evaluation guidance about final files and persistent services. They used a local Docker runtime. The published Kimi K3 field used a shared setup with prepared VM checkpoints and did not receive Graff’s append.

Grok also differs from the published field’s model. The recorded outcomes describe those complete configurations; they do not isolate the effect of switching harnesses alone.

How are cost and time calculated?

Cost per pass is recorded spend divided by completed tasks, including spending on failures. Graff uses historical evaluator list-price estimates. The published field uses cost_first_cold_usd, which reprices first-turn cache reads. These are different accounting methods, not current provider price quotes.

Where task costs are missing, ≥ denotes a lower bound. No passes means the ratio is undefined, not free. Median attempt time includes successful and failed attempts where timings are available. Graff medians are not reported; summary maxima are not medians.

Can I inspect the underlying evidence?

These are repository-reported results, not a new rerun or independent audit. Graff records are pinned to b4bd80c; the published board is pinned to 8f11b13. Publication and page-update dates do not imply new evaluation runs.

THE FULL WRITE-UP

What happened inside the runs?

Task failures, model differences, pricing methods, and the reasoning behind the comparisons.

Read the experiment and methodology