{"schemaVersion":1,"canonical":"https://codegraff.com/blog/graff-frontier-harness-evals","markdown":"https://codegraff.com/blog/graff-frontier-harness-evals/markdown","slug":"graff-frontier-harness-evals","title":"Graff on FrontierHarness: coding agent benchmarks, costs and comparisons","description":"Compare Graff’s recorded runs with the FrontierHarness field: Exo, Codex, Claude Code, Pi and more. Explore task completion, cost per pass and DeepSWE results.","publishedAt":"2026-09-06","intro":"Graff completed 20 of 21 terminal tasks with Grok 4.6 and 17 with Kimi K3 in our later FrontierHarness runs. The public benchmark tests nine coding harnesses in twelve configurations, including Exo, Codex, Claude Code and Pi. Here is what each setup completed, what it cost, and where the comparisons hold.","hero":{"src":"/blog/graff-frontier-harness-evals/eval-workshop.webp","alt":"Three mouse engineers test completed pieces from separate brass workbenches using one central checking jig.","caption":"The same checking jig is only part of a fair comparison. Models, instructions, and execution environments matter too. Conceptual illustration; scores appear in the charts and tables below."},"sourceNote":"Results are repository-reported evaluations checked against the tracked summaries and protocol at revision b4bd80c, not a fresh rerun or an independent audit. The Graff terminal runs shown here used additional evaluation instructions. Models and execution environments differ across some rows. Published Exo costs are taken from FrontierHarness revision 8f11b13; the terminal subtotal is calculated from its public task-level records.","sections":[{"id":"suite","title":"Thirty tasks, two different tests of a harness","paragraphs":["This coding harness comparison focuses on measured task completion. A coding harness connects a model to tools, carries context between steps, and manages execution. Evaluating the model–harness combination tells us whether that system can finish the work, not just describe a solution.","FrontierHarness combines 21 Terminal-Bench 2.1 tasks with nine DeepSWE tasks. The terminal slice grades the final environment with public tests. The DeepSWE slice grades a submitted patch against the real project's verifier. Passing a terminal task and producing a patch that survives a regression suite are different achievements.","We ran Graff with Grok 4.6 and Kimi K3, recorded an Exo run with Grok 4.6, and compared those observations with the published Kimi K3 board. The board spans twelve configurations and thirty tasks. That makes it useful context, but it does not turn every row into a controlled harness-only comparison.","Exo's published 16/30 consists of 16/21 terminal tasks and 0/9 DeepSWE tasks. To make the comparison useful, we compare terminal results with terminal results and give the nine-task patch-verification slice its own table."],"sources":["protocol","results"]},{"id":"terminal","title":"Graff completes 20 of 21 terminal tasks","paragraphs":["Graff/Grok 4.6 passed 20/21 tasks, or 95.2% of the selected terminal slice. Graff/Kimi K3 passed 17/21, or 81.0%. Both are the later evaluation configurations.","The published Exo/Kimi K3 board result is 16/21 terminal passes. Our separately recorded Exo/Grok 4.6 run reports 9/21. Graff therefore has the higher recorded terminal count in these rows, but the table describes configurations rather than isolating the harness as the only changing variable.","Both Graff terminal configurations used BENCH_APPEND: extra evaluation instructions supplied by the runner, including guidance on final files and persistent services. The published K3 board did not receive that append. Our Docker setup also differs from the board’s prepared VM checkpoint. The scores are useful evidence about these runs; they are not a controlled claim that switching harnesses alone produces the difference."],"table":{"caption":"Terminal-Bench slice: 21 tasks. These rows do not all share one protocol.","headers":["Harness / model","Passed","Evaluation condition"],"rows":[["Graff / Grok 4.6","20 / 21 (95.2%)","Recorded later configuration with evaluation instructions."],["Graff / Kimi K3","17 / 21 (81.0%)","Evaluation instructions included; model matches the published K3 board."],["Exo / Grok 4.6, local run","9 / 21 recorded","A different model from the published Exo/K3 board result."],["Exo / Kimi K3, published board","16 / 21","Published board configuration; not given Graff's evaluation append."]]},"sources":["results","protocol","grok-summary","kimi-summary"],"chart":{"kind":"terminal","number":"01","title":"20 of 21 tasks, completed.","note":"Each block represents one task in the 21-task terminal slice. Filled blocks count passes; their positions do not identify matching tasks across runs. Graff used extra evaluation instructions. The published Exo board used different instructions and a different runtime."}},{"id":"field","title":"Graff alongside the full FrontierHarness field","paragraphs":["FrontierHarness holds Kimi K3 and the execution setup constant across its published field. The nine harnesses are Codex, Claude Code, DeepSeek Harness, Pi, Oh My Pi, Kimi Code, Exo Harness, OpenCode and Hermes. DeepSeek Harness appears in four modes: Creator, Standard, Minimal and PTC.","On the full 30-task suite, Codex completes 20 tasks. DSH Creator and Claude Code each complete 19. Pi, DSH Standard and DSH PTC complete 18; Oh My Pi, Kimi Code and DSH Minimal complete 17. Exo finishes 16, followed by OpenCode and Hermes at 15. Those totals combine terminal work and patch verification.","A cost-versus-completion Pareto frontier marks the tradeoffs: no other eligible configuration completes at least as many tasks at no greater cost per pass, with one measure strictly better. Among configurations with complete public cost records, Exo, Pi and Codex form that frontier for the full suite. Missing task costs are flagged below and excluded from frontier badges. DSH Creator, for example, has three tasks without published costs.","The explorer includes both later Graff runs alongside the twelve published configurations. Graff/Grok combines 20 terminal passes and one DeepSWE pass for 21/30; Graff/Kimi K3 combines 17 and one for 18/30. The Graff rows identify their model and extra terminal instructions. Frontier badges still describe only the published Kimi K3 board, because Graff used a different evaluation setup."],"sources":["published-board","board-cost-method","runta-launch","grok-summary","kimi-summary","grok-swe-summary","kimi-swe-summary","results","protocol"],"explorer":{"title":"Find your tradeoff.","note":"Twelve published Kimi K3 configurations plus two recorded Graff runs. Graff used extra terminal instructions and a local Docker runtime; its full-suite totals combine the two evaluated slices. Published costs use cost_first_cold_usd; Graff costs use historical list-price summaries. All recorded failures count toward spend. Partial published costs are lower bounds. Frontier badges consider only fully costed published configurations with at least one pass. Median attempt time includes successes and failures where task timings are available; Graff summary maxima are not medians.","entries":[{"id":"codex","name":"Codex","slices":{"all":{"tasks":30,"passes":20,"totalCost":69.3659265,"costedTasks":30,"costPerPass":3.468296325,"medianSeconds":561.86131},"terminal":{"tasks":21,"passes":15,"totalCost":4.8583788,"costedTasks":21,"costPerPass":0.32389192,"medianSeconds":410.738308},"deepswe":{"tasks":9,"passes":5,"totalCost":64.5075477,"costedTasks":9,"costPerPass":12.90150954,"medianSeconds":2587.068742}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"dsh-creator","name":"DSH Creator","slices":{"all":{"tasks":30,"passes":19,"totalCost":62.4131502,"costedTasks":27,"costPerPass":3.284902642105263,"medianSeconds":748.6974585},"terminal":{"tasks":21,"passes":15,"totalCost":3.3748569,"costedTasks":18,"costPerPass":0.22499046,"medianSeconds":452.85957},"deepswe":{"tasks":9,"passes":4,"totalCost":59.0382933,"costedTasks":9,"costPerPass":14.759573325,"medianSeconds":2232.195525}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"claude-code","name":"Claude Code","slices":{"all":{"tasks":30,"passes":19,"totalCost":348.3998442,"costedTasks":29,"costPerPass":18.33683390526316,"medianSeconds":1108.524014},"terminal":{"tasks":21,"passes":16,"totalCost":15.5788197,"costedTasks":21,"costPerPass":0.97367623125,"medianSeconds":577.971591},"deepswe":{"tasks":9,"passes":3,"totalCost":332.8210245,"costedTasks":8,"costPerPass":110.9403415,"medianSeconds":3595.917111}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"pi-responses","name":"Pi","slices":{"all":{"tasks":30,"passes":18,"totalCost":43.793379,"costedTasks":30,"costPerPass":2.4329655,"medianSeconds":638.7449819999999},"terminal":{"tasks":21,"passes":16,"totalCost":5.4599451,"costedTasks":21,"costPerPass":0.34124656875,"medianSeconds":485.996955},"deepswe":{"tasks":9,"passes":2,"totalCost":38.3334339,"costedTasks":9,"costPerPass":19.16671695,"medianSeconds":2350.514284}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"dsh-ptc","name":"DSH PTC","slices":{"all":{"tasks":30,"passes":18,"totalCost":82.39405289999999,"costedTasks":27,"costPerPass":4.5774473833333325,"medianSeconds":534.2175105},"terminal":{"tasks":21,"passes":16,"totalCost":3.0197976,"costedTasks":18,"costPerPass":0.18873735,"medianSeconds":474.941809},"deepswe":{"tasks":9,"passes":2,"totalCost":79.37425529999999,"costedTasks":9,"costPerPass":39.687127649999994,"medianSeconds":3020.700291}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"dsh-standard","name":"DSH Standard","slices":{"all":{"tasks":30,"passes":18,"totalCost":62.259954,"costedTasks":27,"costPerPass":3.4588863333333335,"medianSeconds":592.9845035000001},"terminal":{"tasks":21,"passes":16,"totalCost":8.1456882,"costedTasks":19,"costPerPass":0.5091055125,"medianSeconds":424.040142},"deepswe":{"tasks":9,"passes":2,"totalCost":54.1142658,"costedTasks":8,"costPerPass":27.0571329,"medianSeconds":2835.545963}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"oh-my-pi","name":"Oh My Pi","slices":{"all":{"tasks":30,"passes":17,"totalCost":80.674218,"costedTasks":30,"costPerPass":4.745542235294117,"medianSeconds":595.9183135000001},"terminal":{"tasks":21,"passes":16,"totalCost":5.4427851,"costedTasks":21,"costPerPass":0.34017406875,"medianSeconds":444.418426},"deepswe":{"tasks":9,"passes":1,"totalCost":75.2314329,"costedTasks":9,"costPerPass":75.2314329,"medianSeconds":2997.462833}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"kimi-code","name":"Kimi Code","slices":{"all":{"tasks":30,"passes":17,"totalCost":62.0004801,"costedTasks":30,"costPerPass":3.6470870647058824,"medianSeconds":560.202486},"terminal":{"tasks":21,"passes":14,"totalCost":9.3917685,"costedTasks":21,"costPerPass":0.6708406071428571,"medianSeconds":475.880631},"deepswe":{"tasks":9,"passes":3,"totalCost":52.6087116,"costedTasks":9,"costPerPass":17.5362372,"medianSeconds":2668.571971}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"dsh-minimal","name":"DSH Minimal","slices":{"all":{"tasks":30,"passes":17,"totalCost":80.2052463,"costedTasks":26,"costPerPass":4.717955664705882,"medianSeconds":893.5280545},"terminal":{"tasks":21,"passes":15,"totalCost":6.2994639,"costedTasks":17,"costPerPass":0.41996426,"medianSeconds":414.517787},"deepswe":{"tasks":9,"passes":2,"totalCost":73.90578239999999,"costedTasks":9,"costPerPass":36.9528912,"medianSeconds":2485.273986}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"exo","name":"Exo Harness","slices":{"all":{"tasks":30,"passes":16,"totalCost":16.7227377,"costedTasks":30,"costPerPass":1.04517110625,"medianSeconds":603.5377495},"terminal":{"tasks":21,"passes":16,"totalCost":6.4250046,"costedTasks":21,"costPerPass":0.4015627875,"medianSeconds":411.090475},"deepswe":{"tasks":9,"passes":0,"totalCost":10.2977331,"costedTasks":9,"costPerPass":null,"medianSeconds":954.5815}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"opencode","name":"OpenCode","slices":{"all":{"tasks":30,"passes":15,"totalCost":48.6651828,"costedTasks":29,"costPerPass":3.24434552,"medianSeconds":701.358377},"terminal":{"tasks":21,"passes":15,"totalCost":3.4770654,"costedTasks":21,"costPerPass":0.23180436,"medianSeconds":463.111223},"deepswe":{"tasks":9,"passes":0,"totalCost":45.1881174,"costedTasks":8,"costPerPass":null,"medianSeconds":2944.994884}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"hermes","name":"Hermes","slices":{"all":{"tasks":30,"passes":15,"totalCost":43.5645522,"costedTasks":28,"costPerPass":2.9043034800000003,"medianSeconds":785.6614105000001},"terminal":{"tasks":21,"passes":13,"totalCost":11.3271525,"costedTasks":21,"costPerPass":0.871319423076923,"medianSeconds":547.60882},"deepswe":{"tasks":9,"passes":2,"totalCost":32.2373997,"costedTasks":7,"costPerPass":16.11869985,"medianSeconds":1728.459216}},"cohort":"published","condition":"Published Kimi K3 board; shared runtime and instructions."},{"id":"graff-grok","name":"Graff / Grok 4.6","cohort":"graff","condition":"Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries.","slices":{"all":{"tasks":30,"passes":21,"totalCost":31.7951,"costedTasks":30,"costPerPass":1.514052380952381,"medianSeconds":null},"terminal":{"tasks":21,"passes":20,"totalCost":6.2335,"costedTasks":21,"costPerPass":0.31167500000000004,"medianSeconds":null},"deepswe":{"tasks":9,"passes":1,"totalCost":25.5616,"costedTasks":9,"costPerPass":25.5616,"medianSeconds":null}}},{"id":"graff-kimi","name":"Graff / Kimi K3","cohort":"graff","condition":"Recorded Graff evaluation; extra terminal instructions and local Docker runtime. Full-suite totals combine the terminal and DeepSWE slices. Historical evaluator list prices. Excluded from the published-board frontier. Median attempt time is not reported in the stored summaries.","slices":{"all":{"tasks":30,"passes":18,"totalCost":25.561148,"costedTasks":30,"costPerPass":1.4200637777777778,"medianSeconds":null},"terminal":{"tasks":21,"passes":17,"totalCost":3.6463,"costedTasks":21,"costPerPass":0.21448823529411765,"medianSeconds":null},"deepswe":{"tasks":9,"passes":1,"totalCost":21.914848,"costedTasks":9,"costPerPass":21.914848,"medianSeconds":null}}}]},"figure":{"src":"/blog/graff-frontier-harness-evals/harness-exhibition.webp","alt":"Mouse engineers compare differently shaped brass coding machines that receive matching task cards.","caption":"Different machines, the same task cards. A conceptual illustration of the published harness comparison."}},{"id":"misses","title":"What a completed terminal task really requires","paragraphs":["Terminal tasks test more than whether a model can produce a plausible code block. The harness must inspect the environment, make the changes, run tools, and leave a result the external tests can use. That final state is where small execution details become decisive.","The evaluation notes highlight examples worth checking in any coding harness: a directory that must contain only the requested source file, a service that must remain available to the verifier, and recovered data that must be complete rather than merely parseable. These are concrete completion requirements, not stylistic preferences.","The extra evaluation instructions make those expectations explicit for Graff. That is part of the configuration behind the showcased scores. A separate held-out evaluation would be needed to establish how much these lessons generalize to tasks that did not inform the instructions."],"sources":["failures","protocol"]},{"id":"deepswe","title":"DeepSWE is the harder boundary: one pass out of nine","paragraphs":["Both Graff runs passed the KaTeX multicolumn-array-spans task and failed the other eight. The repository reports a mix of patch-application failures and genuine test failures. The Grok httpx attempt came close on the fail-to-pass metric, but a near-pass is still a failure under the verifier.","The runner does not ask a model whether a patch looks plausible. It starts the task image, installs the task's test bundle, supplies the patch, runs the prepare and test steps, and reads reward.json. When the verifier result is missing, the run fails rather than receiving an inferred success.","Patch application is part of the result. A patch that applies to the raw working tree can fail against the grader's prepared base. Some cleaned submissions contained no real project change after scratch files were removed. Those are useful failures to expose, not reasons to quietly remove tasks from the denominator."],"table":{"caption":"DeepSWE slice: nine tasks, graded separately from the terminal tasks.","headers":["Harness / model","Passed","Reported outcome"],"rows":[["Graff / Grok 4.6","1 / 9","KaTeX passed; the other eight did not."],["Graff / Kimi K3","1 / 9","KaTeX passed; the other eight did not."],["Exo / Kimi K3, published board","0 / 9","Board result; model and run conditions must stay attached to the comparison."]]},"sources":["results","protocol","grader"],"chart":{"kind":"deepswe","number":"03","title":"Nine tasks. One verified patch per Graff run.","note":"Each circle represents one task in the nine-task DeepSWE slice. Both Graff runs passed KaTeX; eight tasks did not pass. Exo’s published K3 board reports zero passes. These are separate recorded configurations, not repeated trials under identical conditions."}},{"id":"cost","title":"The recorded cost of the terminal runs","paragraphs":["The recorded terminal runs work out to about $0.31 per pass for Graff/Grok, $0.21 for Graff/Kimi K3, and $0.40 for published Exo/Kimi K3. Each ratio includes spending on unsuccessful terminal tasks. The chart compares the same 21-task slice, with the cost basis attached to each row.","Graff’s stored list-price estimates are $6.2335 for Grok and $3.6463 for Kimi K3. For published Exo/Kimi K3, summing cost_first_cold_usd across the 21 terminal-bench task records gives $6.4250046. Dividing by its 16 terminal passes gives $0.4015627875 per pass.","Across the full 30-task suite, published Exo costs $16.7227377 and passes 16 tasks: $1.04517110625 per pass, matching its effective_cost_per_pass field. That is the $1.05 headline in the public README. We keep that full-suite figure out of the terminal-only chart.","The published board reprices first-turn cache reads; Graff uses its evaluator’s historical list-price estimates. Graff also received extra evaluation instructions and ran in a different environment. These recorded ratios do not isolate a harness-only cost advantage. Only our separate local Exo/Grok run lacks token events; published Exo/Kimi costs are available."],"table":{"caption":"Terminal costs across all 21 tasks, including failures. Historical estimates with different accounting and evaluation conditions.","headers":["Run","Passed","Recorded total (USD)","Cost basis"],"rows":[["Graff / Grok 4.6 with evaluation append","20 / 21","$6.2335","Historical evaluator list prices; extra instructions."],["Graff / Kimi K3 with evaluation append","17 / 21","$3.6463","Historical evaluator list prices; extra instructions."],["Exo / Kimi K3, published board","16 / 21","$6.4250046","Sum of published terminal cost_first_cold_usd; first-turn cache repricing."],["Exo / Grok 4.6, local run","9 / 21 recorded","Unavailable: no token events reported","Separate local Grok run; usage not recorded."]]},"sources":["grok-summary","kimi-summary","protocol","published-board","board-cost-method"],"chart":{"kind":"cost","number":"04","title":"The dollars behind each pass.","note":"All three plotted rows cover the 21 terminal tasks. Total cost, including failed attempts, is divided by passes. Exo uses published first-turn cache repricing; Graff uses historical list-price estimates and extra evaluation instructions. Runtime and model conditions differ. The separate local Exo/Grok run has no usage data."},"figure":{"src":"/blog/graff-frontier-harness-evals/cost-accounting.webp","alt":"Two mouse engineers count brass tokens and verify task cards beside a curling receipt and discarded failed attempts.","caption":"The receipt includes the attempts that failed. Cost per pass uses the entire bill for the selected task slice."}},{"id":"meaning","title":"What the results mean when choosing a coding harness","paragraphs":["For people using Graff, the useful implication is that finishing means verifying the requested artifact and its surrounding state. A service must still be running when it is needed. A patch must apply to the expected base. A directory must contain only the files the task allows. These are harness responsibilities as well as model behaviors.","The results support investigating stronger final-state verification and clearer task-boundary handling. They do not establish that the specific evaluation coaching should become a universal system prompt. The failure notes explicitly say that the production prompt file was not changed for this diagnosis.","For a team evaluating coding tools, these results suggest a useful shortlist test: run the same tasks through each candidate with your intended model and instruction configuration. Keep completed tasks as the primary outcome, then compare tokens, time, and memory. Hold the task pack, verifier revision, and execution environment fixed, and repeat the runs before treating a small difference as dependable.","The previous Graff runtime article examined repeated context work. This evaluation asks a different question: did the final result pass the task's verifier? Both are needed. Lower overhead helps only when the task is actually completed."],"sources":["protocol","failures","previous"]},{"id":"discussion","title":"What readers and other articles noticed","paragraphs":["Runta’s launch article centers the gap between task completion and the bill. It highlights Codex for completion, Exo for cost, Pi as a middle ground and DSH Minimal for speed. It also discusses how Claude Code’s model and gateway configuration could affect caching. These are the benchmark publisher’s findings, not independent product reviews.","AgentConn’s follow-up, “Same Model, 17x the Cost: How Your Harness Decides Who Wins”, focuses on the harness cost gap. The underlying cost figures here come from the public evaluation records, so readers can check the arithmetic directly.","In a public r/PiCodingAgent discussion, readers asked for more models and repeated runs before generalizing the ordering. Others questioned how closely the task mix matched everyday development. These are reader reactions, not a representative survey or customer reviews of Graff.","For this comparison, the concrete questions are task completion, cost per successful task, cache behavior and time. A result on one frozen coding harness configuration answers those questions for that run. It does not tell every team which tool to use."],"sources":["runta-launch","agentconn-coverage","reader-discussion"]},{"id":"questions","title":"Questions about this coding agent benchmark","paragraphs":["Which coding harness leads FrontierHarness? On the published 30-task Kimi K3 board, Codex has the most passes at 20/30. Exo has the lowest reported effective cost per pass at about $1.05. Those are different measures; the explorer lets you compare both.","Is this Graff’s score on the full Terminal-Bench leaderboard? No. The 20/21 result covers FrontierHarness’s selected 21-task terminal slice with the documented evaluation instructions. It is not a score across the complete Terminal-Bench task set.","Does Graff beat Exo with the same model? The recorded Kimi K3 rows show 17/21 for Graff and 16/21 for published Exo, but Graff used additional evaluation instructions and a different runtime setup. The counts compare those configurations; they do not isolate a harness-only advantage.","What does cost per completed task include? All spending on the selected slice, including unsuccessful attempts, divided by passes. Published Exo/Kimi K3 is about $0.40 per pass for the 21 terminal tasks, versus $1.05 for all 30 tasks. The chart uses the terminal slice; its accounting differs from Graff’s historical list-price estimates. Only the separate local Exo/Grok run has missing usage.","Can an agent inspect the underlying evidence? This article has Markdown and structured JSON representations containing the same paragraphs, result tables, and source URLs as the page. Each result links to the pinned evaluation records."],"sources":["protocol","results","grok-summary","kimi-summary","published-board","board-cost-method","runta-launch"]},{"id":"inspect","title":"Explore the evaluation and try Graff","paragraphs":["The evaluation directory includes the runner, protocol, result summaries, and DeepSWE grading code. Start with PROTOCOL.md before reproducing a headline score: the runner's evaluation append, container preparation, and model choice affect what the result means.","The local runs used Docker task images rather than the board's prepared VM checkpoint. On stripped images, the runner added the certificate bundle and test dependencies needed to execute the work. That may be reasonable evaluation setup, but it is another reason not to describe the environments as identical.","The later terminal results are the highlight: 20/21 with Grok 4.6 and 17/21 with Kimi K3, under the documented evaluation configuration. The separate DeepSWE results show the remaining challenge in producing patches that pass full verification. Together, they offer a more useful picture of Graff than a score stripped of its task and model."],"sources":["protocol","results"]}],"sources":[{"id":"results","title":"FrontierHarness results and suite definitions","url":"https://github.com/justrach/codegraff/blob/b4bd80c0dbf4332f249b4a1a36fbeb136ed65611/graff-evals/frontier-harness/README.md"},{"id":"protocol","title":"Evaluation protocol and comparison limitations","url":"https://github.com/justrach/codegraff/blob/b4bd80c0dbf4332f249b4a1a36fbeb136ed65611/graff-evals/frontier-harness/PROTOCOL.md"},{"id":"grok-summary","title":"Graff / Grok terminal result summary","url":"https://github.com/justrach/codegraff/blob/b4bd80c0dbf4332f249b4a1a36fbeb136ed65611/graff-evals/frontier-harness/results.maxima.json"},{"id":"kimi-summary","title":"Graff / Kimi K3 terminal result summary","url":"https://github.com/justrach/codegraff/blob/b4bd80c0dbf4332f249b4a1a36fbeb136ed65611/graff-evals/frontier-harness/kimi-results.maxima.json"},{"id":"failures","title":"Terminal task failure diagnoses","url":"https://github.com/justrach/codegraff/blob/b4bd80c0dbf4332f249b4a1a36fbeb136ed65611/graff-evals/frontier-harness/FAILURES.md"},{"id":"grader","title":"DeepSWE patch verifier integration","url":"https://github.com/justrach/codegraff/blob/b4bd80c0dbf4332f249b4a1a36fbeb136ed65611/graff-evals/frontier-harness/grade_swe.py"},{"id":"previous","title":"Graff’s RLM runtime: how 135k input tokens became 25.3k","url":"https://codegraff.com/blog/graff-rlm-token-efficiency"},{"id":"published-board","title":"Published FrontierHarness task costs and Exo results","url":"https://github.com/frontier-harness-eval/eval/blob/8f11b130c30bbf76ca1f3edeea70abc773bd8d2c/results/eval-data.json"},{"id":"board-cost-method","title":"Published cost accounting and comparison method","url":"https://github.com/frontier-harness-eval/eval/blob/8f11b130c30bbf76ca1f3edeea70abc773bd8d2c/skills/frontierharness-eval/reference.md#cost-comparability"},{"id":"runta-launch","title":"Runta: Introducing FrontierHarness Eval","url":"https://runta.com/blog/introducing-frontierharness-eval/"},{"id":"agentconn-coverage","title":"AgentConn: Same Model, 17x the Cost","url":"https://agentconn.com/blog/eval-harness-benchmark-cost-2026/"},{"id":"reader-discussion","title":"Reader discussion of FrontierHarness in r/PiCodingAgent","url":"https://www.reddit.com/r/PiCodingAgent/comments/1w5okw5/frontierharness_eval_looks_specifically_at_coding/"},{"id":"grok-swe-summary","title":"Graff / Grok DeepSWE cost summary","url":"https://github.com/justrach/codegraff/blob/b4bd80c0dbf4332f249b4a1a36fbeb136ed65611/graff-evals/frontier-harness/swe-results.maxima.json"},{"id":"kimi-swe-summary","title":"Graff / Kimi K3 DeepSWE cost summary","url":"https://github.com/justrach/codegraff/blob/b4bd80c0dbf4332f249b4a1a36fbeb136ed65611/graff-evals/frontier-harness/swe-kimi-results.maxima.json"}],"seoTitle":"Graff vs Exo, Codex & Claude Code | Harness Benchmarks","keywords":["coding harness comparison","AI coding agent benchmarks","FrontierHarness Eval","Graff vs Exo","Codex vs Claude Code benchmarks","Pi coding agent benchmark","DeepSeek Harness","Kimi K3 harness comparison","coding agent cost per pass","coding agent Pareto frontier","Terminal-Bench","DeepSWE","prompt caching"]}