Verify code changes
Check a repository change
- Read the repository instructions and reproduce the specific behavior being changed. Record the starting revision and any existing failures.
- Make a focused patch. Add a regression test when it checks meaningful behavior; use the project’s documented test runner.
- Ask Graff to execute the relevant tests and required build or lint checks. Inspect command output and exit status. If a check fails, determine whether the patch caused it, repair the issue, and rerun the affected check.
- Review the final diff for unrelated edits, missing cases and behavior the tests do not cover. Keep skipped, blocked and failing checks visible.
This is an explicit acceptance workflow. Do not assume every edit automatically runs tests, or that a successful shell command covers every requirement. Start with the worked CLI example.
Keep an evidence record
For each reviewed change, retain the base revision, patch, Graff version, selected model and route, environment, exact check commands, exit statuses and relevant raw output. Redact credentials and private data before sharing a record.
Base revision: <commit> Patch: <reviewed diff or commit> Graff version / model / route: <recorded values> Environment: <OS, runtime, dependency lockfile> Check: <exact command> Exit status: <recorded status, or unavailable> Output: <path to retained output> Outcome: <passed, failed, skipped or blocked> Limitations: <behaviors not checked>
Use graff --version and git rev-parse HEAD for the version and starting commit. Save the final patch with git diff; include newly added files in your review because an unstaged diff does not show untracked file contents.
Missing output means the check is unverified. Passing tests support the behaviors they cover; they do not establish that every possible input or environment works.
Read saved evaluations
The Graff evaluations page reports selected terminal tasks and DeepSWE patch tasks. It links the original experiment and source records. The JSON data provides the published aggregate comparison.
Read a configuration together with its task slice, model, instructions, runtime and cost method. Follow the linked source summary for the recorded task outcomes and any available artifacts; the comparison page is not a complete raw session log.
- The 21 selected terminal tasks and nine DeepSWE patch tasks measure different outcomes. These are selected slices, not full leaderboard scores.
- The Graff terminal runs used additional instructions and a local Docker runtime. Comparisons with the published field do not isolate a harness-only effect.
- Missing verifier results count as failures in this benchmark. That scoring rule does not turn missing measurements in other reports into observed failures.
- Recorded costs include failed attempts. Partial costs are lower bounds; absent costs or timings are unknown.
For a specific example, the published DeepSWE summaries record a pass on the KaTeX multicolumn-array-spans task for both Graff configurations; the other eight tasks did not pass. Inspect the source summaries linked from evaluation methodology before generalizing to your repository.