Graff completed 20 of 21 terminal tasks with Grok 4.6 and 17 with Kimi K3 in our later FrontierHarness runs. The public benchmark tests nine coding harnesses in twelve configurations, including Exo, Codex, Claude Code and Pi. Here is what each setup completed, what it cost, and where the comparisons hold.

Thirty tasks, two different tests of a harness
This coding harness comparison focuses on measured task completion. A coding harness connects a model to tools, carries context between steps, and manages execution. Evaluating the model–harness combination tells us whether that system can finish the work, not just describe a solution.
FrontierHarness combines 21 Terminal-Bench 2.1 tasks with nine DeepSWE tasks. The terminal slice grades the final environment with public tests. The DeepSWE slice grades a submitted patch against the real project's verifier. Passing a terminal task and producing a patch that survives a regression suite are different achievements.
We ran Graff with Grok 4.6 and Kimi K3, recorded an Exo run with Grok 4.6, and compared those observations with the published Kimi K3 board. The board spans twelve configurations and thirty tasks. That makes it useful context, but it does not turn every row into a controlled harness-only comparison.
Exo's published 16/30 consists of 16/21 terminal tasks and 0/9 DeepSWE tasks. To make the comparison useful, we compare terminal results with terminal results and give the nine-task patch-verification slice its own table.
Sources: Evaluation protocol and comparison limitations · FrontierHarness results and suite definitions
Graff completes 20 of 21 terminal tasks
Graff/Grok 4.6 passed 20/21 tasks, or 95.2% of the selected terminal slice. Graff/Kimi K3 passed 17/21, or 81.0%. Both are the later evaluation configurations.
20 of 21 tasks, completed.
The published Exo/Kimi K3 board result is 16/21 terminal passes. Our separately recorded Exo/Grok 4.6 run reports 9/21. Graff therefore has the higher recorded terminal count in these rows, but the table describes configurations rather than isolating the harness as the only changing variable.
Both Graff terminal configurations used BENCH_APPEND: extra evaluation instructions supplied by the runner, including guidance on final files and persistent services. The published K3 board did not receive that append. Our Docker setup also differs from the board’s prepared VM checkpoint. The scores are useful evidence about these runs; they are not a controlled claim that switching harnesses alone produces the difference.
Inspect data & evaluation conditions
| Harness / model | Passed | Evaluation condition |
|---|---|---|
| Graff / Grok 4.6 | 20 / 21 (95.2%) | Recorded later configuration with evaluation instructions. |
| Graff / Kimi K3 | 17 / 21 (81.0%) | Evaluation instructions included; model matches the published K3 board. |
| Exo / Grok 4.6, local run | 9 / 21 recorded | A different model from the published Exo/K3 board result. |
| Exo / Kimi K3, published board | 16 / 21 | Published board configuration; not given Graff's evaluation append. |
Sources: FrontierHarness results and suite definitions · Evaluation protocol and comparison limitations · Graff / Grok terminal result summary · Graff / Kimi K3 terminal result summary
Graff alongside the full FrontierHarness field
FrontierHarness holds Kimi K3 and the execution setup constant across its published field. The nine harnesses are Codex, Claude Code, DeepSeek Harness, Pi, Oh My Pi, Kimi Code, Exo Harness, OpenCode and Hermes. DeepSeek Harness appears in four modes: Creator, Standard, Minimal and PTC.
On the full 30-task suite, Codex completes 20 tasks. DSH Creator and Claude Code each complete 19. Pi, DSH Standard and DSH PTC complete 18; Oh My Pi, Kimi Code and DSH Minimal complete 17. Exo finishes 16, followed by OpenCode and Hermes at 15. Those totals combine terminal work and patch verification.
A cost-versus-completion Pareto frontier marks the tradeoffs: no other eligible configuration completes at least as many tasks at no greater cost per pass, with one measure strictly better. Among configurations with complete public cost records, Exo, Pi and Codex form that frontier for the full suite. Missing task costs are flagged below and excluded from frontier badges. DSH Creator, for example, has three tasks without published costs.
The explorer includes both later Graff runs alongside the twelve published configurations. Graff/Grok combines 20 terminal passes and one DeepSWE pass for 21/30; Graff/Kimi K3 combines 17 and one for 18/30. The Graff rows identify their model and extra terminal instructions. Frontier badges still describe only the published Kimi K3 board, because Graff used a different evaluation setup.
Find your tradeoff.
Two Graff runs. Twelve published configurations. Explore them together.
14 configurations · All 30 tasks · More completed is better
- Completed
- 21/30
- Cost per pass
- $1.51
- Median attempt
- Not reported
Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries. 30/30 tasks covered by recorded costs. Recorded spend: $31.7951. Failed attempts are included.
Inspect all 42 configuration and slice records
| Configuration | Passes | Recorded total | Cost / pass | Costed tasks | Median attempt time |
|---|---|---|---|---|---|
| Codex Published Kimi K3 board; shared runtime and instructions. | 20/30 | $69.3659 | $3.47 | 30/30 | 9m 22s |
| DSH Creator Published Kimi K3 board; shared runtime and instructions. | 19/30 | ≥ $62.4132 | ≥ $3.28 | 27/30 | 12m 29s |
| Claude Code Published Kimi K3 board; shared runtime and instructions. | 19/30 | ≥ $348.3998 | ≥ $18.34 | 29/30 | 18m 29s |
| Pi Published Kimi K3 board; shared runtime and instructions. | 18/30 | $43.7934 | $2.43 | 30/30 | 10m 39s |
| DSH PTC Published Kimi K3 board; shared runtime and instructions. | 18/30 | ≥ $82.3941 | ≥ $4.58 | 27/30 | 8m 54s |
| DSH Standard Published Kimi K3 board; shared runtime and instructions. | 18/30 | ≥ $62.2600 | ≥ $3.46 | 27/30 | 9m 53s |
| Oh My Pi Published Kimi K3 board; shared runtime and instructions. | 17/30 | $80.6742 | $4.75 | 30/30 | 9m 56s |
| Kimi Code Published Kimi K3 board; shared runtime and instructions. | 17/30 | $62.0005 | $3.65 | 30/30 | 9m 20s |
| DSH Minimal Published Kimi K3 board; shared runtime and instructions. | 17/30 | ≥ $80.2052 | ≥ $4.72 | 26/30 | 14m 54s |
| Exo Harness Published Kimi K3 board; shared runtime and instructions. | 16/30 | $16.7227 | $1.05 | 30/30 | 10m 04s |
| OpenCode Published Kimi K3 board; shared runtime and instructions. | 15/30 | ≥ $48.6652 | ≥ $3.24 | 29/30 | 11m 41s |
| Hermes Published Kimi K3 board; shared runtime and instructions. | 15/30 | ≥ $43.5646 | ≥ $2.90 | 28/30 | 13m 06s |
| Graff / Grok 4.6 Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries. | 21/30 | $31.7951 | $1.51 | 30/30 | Not reported |
| Graff / Kimi K3 Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries. | 18/30 | $25.5611 | $1.42 | 30/30 | Not reported |
| Configuration | Passes | Recorded total | Cost / pass | Costed tasks | Median attempt time |
|---|---|---|---|---|---|
| Codex Published Kimi K3 board; shared runtime and instructions. | 15/21 | $4.8584 | $0.32 | 21/21 | 6m 51s |
| DSH Creator Published Kimi K3 board; shared runtime and instructions. | 15/21 | ≥ $3.3749 | ≥ $0.22 | 18/21 | 7m 33s |
| Claude Code Published Kimi K3 board; shared runtime and instructions. | 16/21 | $15.5788 | $0.97 | 21/21 | 9m 38s |
| Pi Published Kimi K3 board; shared runtime and instructions. | 16/21 | $5.4599 | $0.34 | 21/21 | 8m 06s |
| DSH PTC Published Kimi K3 board; shared runtime and instructions. | 16/21 | ≥ $3.0198 | ≥ $0.19 | 18/21 | 7m 55s |
| DSH Standard Published Kimi K3 board; shared runtime and instructions. | 16/21 | ≥ $8.1457 | ≥ $0.51 | 19/21 | 7m 04s |
| Oh My Pi Published Kimi K3 board; shared runtime and instructions. | 16/21 | $5.4428 | $0.34 | 21/21 | 7m 24s |
| Kimi Code Published Kimi K3 board; shared runtime and instructions. | 14/21 | $9.3918 | $0.67 | 21/21 | 7m 56s |
| DSH Minimal Published Kimi K3 board; shared runtime and instructions. | 15/21 | ≥ $6.2995 | ≥ $0.42 | 17/21 | 6m 55s |
| Exo Harness Published Kimi K3 board; shared runtime and instructions. | 16/21 | $6.4250 | $0.40 | 21/21 | 6m 51s |
| OpenCode Published Kimi K3 board; shared runtime and instructions. | 15/21 | $3.4771 | $0.23 | 21/21 | 7m 43s |
| Hermes Published Kimi K3 board; shared runtime and instructions. | 13/21 | $11.3272 | $0.87 | 21/21 | 9m 08s |
| Graff / Grok 4.6 Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries. | 20/21 | $6.2335 | $0.31 | 21/21 | Not reported |
| Graff / Kimi K3 Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries. | 17/21 | $3.6463 | $0.21 | 21/21 | Not reported |
| Configuration | Passes | Recorded total | Cost / pass | Costed tasks | Median attempt time |
|---|---|---|---|---|---|
| Codex Published Kimi K3 board; shared runtime and instructions. | 5/9 | $64.5075 | $12.90 | 9/9 | 43m 07s |
| DSH Creator Published Kimi K3 board; shared runtime and instructions. | 4/9 | $59.0383 | $14.76 | 9/9 | 37m 12s |
| Claude Code Published Kimi K3 board; shared runtime and instructions. | 3/9 | ≥ $332.8210 | ≥ $110.94 | 8/9 | 59m 56s |
| Pi Published Kimi K3 board; shared runtime and instructions. | 2/9 | $38.3334 | $19.17 | 9/9 | 39m 11s |
| DSH PTC Published Kimi K3 board; shared runtime and instructions. | 2/9 | $79.3743 | $39.69 | 9/9 | 50m 21s |
| DSH Standard Published Kimi K3 board; shared runtime and instructions. | 2/9 | ≥ $54.1143 | ≥ $27.06 | 8/9 | 47m 16s |
| Oh My Pi Published Kimi K3 board; shared runtime and instructions. | 1/9 | $75.2314 | $75.23 | 9/9 | 49m 57s |
| Kimi Code Published Kimi K3 board; shared runtime and instructions. | 3/9 | $52.6087 | $17.54 | 9/9 | 44m 29s |
| DSH Minimal Published Kimi K3 board; shared runtime and instructions. | 2/9 | $73.9058 | $36.95 | 9/9 | 41m 25s |
| Exo Harness Published Kimi K3 board; shared runtime and instructions. | 0/9 | $10.2977 | No passes | 9/9 | 15m 55s |
| OpenCode Published Kimi K3 board; shared runtime and instructions. | 0/9 | ≥ $45.1881 | No passes | 8/9 | 49m 05s |
| Hermes Published Kimi K3 board; shared runtime and instructions. | 2/9 | ≥ $32.2374 | ≥ $16.12 | 7/9 | 28m 48s |
| Graff / Grok 4.6 Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries. | 1/9 | $25.5616 | $25.56 | 9/9 | Not reported |
| Graff / Kimi K3 Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries. | 1/9 | $21.9148 | $21.91 | 9/9 | Not reported |

Sources: Published FrontierHarness task costs and Exo results · Published cost accounting and comparison method · Runta: Introducing FrontierHarness Eval · Graff / Grok terminal result summary · Graff / Kimi K3 terminal result summary · Graff / Grok DeepSWE cost summary · Graff / Kimi K3 DeepSWE cost summary · FrontierHarness results and suite definitions · Evaluation protocol and comparison limitations
What a completed terminal task really requires
Terminal tasks test more than whether a model can produce a plausible code block. The harness must inspect the environment, make the changes, run tools, and leave a result the external tests can use. That final state is where small execution details become decisive.
The evaluation notes highlight examples worth checking in any coding harness: a directory that must contain only the requested source file, a service that must remain available to the verifier, and recovered data that must be complete rather than merely parseable. These are concrete completion requirements, not stylistic preferences.
The extra evaluation instructions make those expectations explicit for Graff. That is part of the configuration behind the showcased scores. A separate held-out evaluation would be needed to establish how much these lessons generalize to tasks that did not inform the instructions.
Sources: Terminal task failure diagnoses · Evaluation protocol and comparison limitations
DeepSWE is the harder boundary: one pass out of nine
Both Graff runs passed the KaTeX multicolumn-array-spans task and failed the other eight. The repository reports a mix of patch-application failures and genuine test failures. The Grok httpx attempt came close on the fail-to-pass metric, but a near-pass is still a failure under the verifier.
Nine tasks. One verified patch per Graff run.
The runner does not ask a model whether a patch looks plausible. It starts the task image, installs the task's test bundle, supplies the patch, runs the prepare and test steps, and reads reward.json. When the verifier result is missing, the run fails rather than receiving an inferred success.
Patch application is part of the result. A patch that applies to the raw working tree can fail against the grader's prepared base. Some cleaned submissions contained no real project change after scratch files were removed. Those are useful failures to expose, not reasons to quietly remove tasks from the denominator.
Inspect data & evaluation conditions
| Harness / model | Passed | Reported outcome |
|---|---|---|
| Graff / Grok 4.6 | 1 / 9 | KaTeX passed; the other eight did not. |
| Graff / Kimi K3 | 1 / 9 | KaTeX passed; the other eight did not. |
| Exo / Kimi K3, published board | 0 / 9 | Board result; model and run conditions must stay attached to the comparison. |
Sources: FrontierHarness results and suite definitions · Evaluation protocol and comparison limitations · DeepSWE patch verifier integration
The recorded cost of the terminal runs
The recorded terminal runs work out to about $0.31 per pass for Graff/Grok, $0.21 for Graff/Kimi K3, and $0.40 for published Exo/Kimi K3. Each ratio includes spending on unsuccessful terminal tasks. The chart compares the same 21-task slice, with the cost basis attached to each row.
The dollars behind each pass.
Graff’s stored list-price estimates are $6.2335 for Grok and $3.6463 for Kimi K3. For published Exo/Kimi K3, summing cost_first_cold_usd across the 21 terminal-bench task records gives $6.4250046. Dividing by its 16 terminal passes gives $0.4015627875 per pass.
Across the full 30-task suite, published Exo costs $16.7227377 and passes 16 tasks: $1.04517110625 per pass, matching its effective_cost_per_pass field. That is the $1.05 headline in the public README. We keep that full-suite figure out of the terminal-only chart.
The published board reprices first-turn cache reads; Graff uses its evaluator’s historical list-price estimates. Graff also received extra evaluation instructions and ran in a different environment. These recorded ratios do not isolate a harness-only cost advantage. Only our separate local Exo/Grok run lacks token events; published Exo/Kimi costs are available.

Inspect data & evaluation conditions
| Run | Passed | Recorded total (USD) | Cost basis |
|---|---|---|---|
| Graff / Grok 4.6 with evaluation append | 20 / 21 | $6.2335 | Historical evaluator list prices; extra instructions. |
| Graff / Kimi K3 with evaluation append | 17 / 21 | $3.6463 | Historical evaluator list prices; extra instructions. |
| Exo / Kimi K3, published board | 16 / 21 | $6.4250046 | Sum of published terminal cost_first_cold_usd; first-turn cache repricing. |
| Exo / Grok 4.6, local run | 9 / 21 recorded | Unavailable: no token events reported | Separate local Grok run; usage not recorded. |
Sources: Graff / Grok terminal result summary · Graff / Kimi K3 terminal result summary · Evaluation protocol and comparison limitations · Published FrontierHarness task costs and Exo results · Published cost accounting and comparison method
What the results mean when choosing a coding harness
For people using Graff, the useful implication is that finishing means verifying the requested artifact and its surrounding state. A service must still be running when it is needed. A patch must apply to the expected base. A directory must contain only the files the task allows. These are harness responsibilities as well as model behaviors.
The results support investigating stronger final-state verification and clearer task-boundary handling. They do not establish that the specific evaluation coaching should become a universal system prompt. The failure notes explicitly say that the production prompt file was not changed for this diagnosis.
For a team evaluating coding tools, these results suggest a useful shortlist test: run the same tasks through each candidate with your intended model and instruction configuration. Keep completed tasks as the primary outcome, then compare tokens, time, and memory. Hold the task pack, verifier revision, and execution environment fixed, and repeat the runs before treating a small difference as dependable.
The previous Graff runtime article examined repeated context work. This evaluation asks a different question: did the final result pass the task's verifier? Both are needed. Lower overhead helps only when the task is actually completed.
Sources: Evaluation protocol and comparison limitations · Terminal task failure diagnoses · Graff’s RLM runtime: how 135k input tokens became 25.3k
What readers and other articles noticed
Runta’s launch article centers the gap between task completion and the bill. It highlights Codex for completion, Exo for cost, Pi as a middle ground and DSH Minimal for speed. It also discusses how Claude Code’s model and gateway configuration could affect caching. These are the benchmark publisher’s findings, not independent product reviews.
AgentConn’s follow-up, “Same Model, 17x the Cost: How Your Harness Decides Who Wins”, focuses on the harness cost gap. The underlying cost figures here come from the public evaluation records, so readers can check the arithmetic directly.
In a public r/PiCodingAgent discussion, readers asked for more models and repeated runs before generalizing the ordering. Others questioned how closely the task mix matched everyday development. These are reader reactions, not a representative survey or customer reviews of Graff.
For this comparison, the concrete questions are task completion, cost per successful task, cache behavior and time. A result on one frozen coding harness configuration answers those questions for that run. It does not tell every team which tool to use.
Sources: Runta: Introducing FrontierHarness Eval · AgentConn: Same Model, 17x the Cost · Reader discussion of FrontierHarness in r/PiCodingAgent
Questions about this coding agent benchmark
Which coding harness leads FrontierHarness? On the published 30-task Kimi K3 board, Codex has the most passes at 20/30. Exo has the lowest reported effective cost per pass at about $1.05. Those are different measures; the explorer lets you compare both.
Is this Graff’s score on the full Terminal-Bench leaderboard? No. The 20/21 result covers FrontierHarness’s selected 21-task terminal slice with the documented evaluation instructions. It is not a score across the complete Terminal-Bench task set.
Does Graff beat Exo with the same model? The recorded Kimi K3 rows show 17/21 for Graff and 16/21 for published Exo, but Graff used additional evaluation instructions and a different runtime setup. The counts compare those configurations; they do not isolate a harness-only advantage.
What does cost per completed task include? All spending on the selected slice, including unsuccessful attempts, divided by passes. Published Exo/Kimi K3 is about $0.40 per pass for the 21 terminal tasks, versus $1.05 for all 30 tasks. The chart uses the terminal slice; its accounting differs from Graff’s historical list-price estimates. Only the separate local Exo/Grok run has missing usage.
Can an agent inspect the underlying evidence? This article has Markdown and structured JSON representations containing the same paragraphs, result tables, and source URLs as the page. Each result links to the pinned evaluation records.
Sources: Evaluation protocol and comparison limitations · FrontierHarness results and suite definitions · Graff / Grok terminal result summary · Graff / Kimi K3 terminal result summary · Published FrontierHarness task costs and Exo results · Published cost accounting and comparison method · Runta: Introducing FrontierHarness Eval
Explore the evaluation and try Graff
The evaluation directory includes the runner, protocol, result summaries, and DeepSWE grading code. Start with PROTOCOL.md before reproducing a headline score: the runner's evaluation append, container preparation, and model choice affect what the result means.
The local runs used Docker task images rather than the board's prepared VM checkpoint. On stripped images, the runner added the certificate bundle and test dependencies needed to execute the work. That may be reasonable evaluation setup, but it is another reason not to describe the environments as identical.
The later terminal results are the highlight: 20/21 with Grok 4.6 and 17/21 with Kimi K3, under the documented evaluation configuration. The separate DeepSWE results show the remaining challenge in producing patches that pass full verification. Together, they offer a more useful picture of Graff than a score stripped of its task and model.
Sources: Evaluation protocol and comparison limitations · FrontierHarness results and suite definitions
Sources and method
Results are repository-reported evaluations checked against the tracked summaries and protocol at revision b4bd80c, not a fresh rerun or an independent audit. The Graff terminal runs shown here used additional evaluation instructions. Models and execution environments differ across some rows. Published Exo costs are taken from FrontierHarness revision 8f11b13; the terminal subtotal is calculated from its public task-level records.
- FrontierHarness results and suite definitions
- Evaluation protocol and comparison limitations
- Graff / Grok terminal result summary
- Graff / Kimi K3 terminal result summary
- Terminal task failure diagnoses
- DeepSWE patch verifier integration
- Graff’s RLM runtime: how 135k input tokens became 25.3k
- Published FrontierHarness task costs and Exo results
- Published cost accounting and comparison method
- Runta: Introducing FrontierHarness Eval
- AgentConn: Same Model, 17x the Cost
- Reader discussion of FrontierHarness in r/PiCodingAgent
- Graff / Grok DeepSWE cost summary
- Graff / Kimi K3 DeepSWE cost summary