Skip to article

Field notes / runtime

Graff's RLM runtime: how 135k input tokens became 25.3k

Rach Pradhan · 10 min read · Graff v0.0.276 + v0.0.277

Codegraff workshop mice operate a circular card rail that returns context and evidence to a central model bench.
The runtime is the workshop around the model: context enters, evidence is checked, and useful state returns for another pass.
On this page

The plain-English version

A model is the programmer. The harness is the desk around it.

A wasteful desk makes the programmer reread the same file, carry every tool manual into every task, wait for one job to finish before starting the next, and paste bulky tool results back into every future conversation. A better desk keeps useful working state, starts safe work early, and throws away context that no longer earns its cost.

That is the point of Graff's RLM runtime. The headline result was a stable bind-reuse run on Grok 4.6: 135k to 25.3k input tokens, 99.6 to 39.2 seconds, and 11 to 5 model calls. The result is real for that one run. It is not an 81% promise for every coding task.

TL;DR

  • v0.0.276 made RLM plus mid-stream SpecPTC the default loop.
  • v0.0.277 brought loaded MCP tools into that loop, learned value-free result shapes, and slimmed later results.
  • Grok 4.6 can now use xAI's hosted X search instead of a Graff-run scrape.
  • Every number below is a maintainer-run deployment signal, not an independent leaderboard.

What shipped where

The announcement thread (opens in a new tab) bundled two consecutive releases into one story. The distinction matters.

v0.0.276 (opens in a new tab) shipped on August 25. It made the RLM loop the default and published the bind-reuse, five-task RLM suite, prompt-cache, and grok-build numbers. v0.0.277 (opens in a new tab) followed on August 26 with MCP-inside-RLM, learned result slimming, late RLM showcase, and hosted X search.

So the 135k to 25.3k result belongs to the v0.0.276 stable-reuse case. v0.0.277 made the same prompt-economics argument apply to connected tools and X search.

RLM + SpecPTC: program the context, overlap the wait

Recursive Language Models treat a long prompt as data in an external environment. The model can inspect slices, write code over them, and recursively ask another model about focused fragments. The original RLM paper (opens in a new tab) established the pattern. Graff implements a Zig-native subset around its existing file, shell, code-intelligence, and subagent tools.

Speculative programmatic tool calling adds timing. Once a streamed program contains a closed statement with sufficiently stable inputs, Graff can launch the host call as a future. When the authoritative program reaches the same call, it claims that result or falls back to normal execution. The upstream SpecPTC write-up (opens in a new tab) explains the scheduling idea in detail.

Persistent binds are the less glamorous and more dependable win. A turn can assigns = read_file(path), then use s later without asking the model to rediscover the file and repeat the request path.

Two Codegraff mouse-operated rail loops compare a persistent RLM loop with a speculative branch that rejoins at a verification gate.
RLM keeps state inside a reusable loop. SpecPTC adds an early branch that must rejoin through verification.

Cache discipline: keep the prefix boring

Prompt caching rewards repetition. Change the system prompt or reorder the tool catalog and the reusable prefix can disappear. Graff keeps the catalog head byte-stable and puts loaded schemas in the tool result instead of rewriting the definitions on the next call. xAI's cache guidance (opens in a new tab) supports the stable-prefix and routing approach, although no harness can guarantee a hit.

The Twitter thread reported a 99.8% cached-token ratio in one live Grok 4.6 run. We did not publish the raw response log for that exact denominator, so treat it as an observed run, not a standing guarantee.

v0.0.277 also stopped advertising RLM on every tiny task. It appears when forced, when a batch reaches four native tools, when context reaches half the model's compaction threshold, or when explicitly loaded. The schema costs roughly 200 input tokens before the model writes a program, so small turns should not pay for it.

Codegraff mice route repeated blue-backed context cards through a fast cache path while a changed card returns to a slower press.
A stable prefix lets later passes reuse cached work. Change that prefix and the shortcut disappears.

MCP learned slim: remember shape, not values

Connected tools often return much more JSON than a task needs. If the harness pastes that object into history, every later call can repay for it. v0.0.277 lets loaded MCP names run as RLM host functions, then infers a value-free return shape made of keys and broad types. That shape is stored in .graff/mcp-shapes.json and attached to a later schema load, never the always-on prompt head.

The implementation can slim issue-shaped arrays to identity fields and comment arrays to a count plus latest author. This is deliberately mechanical. An earlier explicit prompt recipe told Grok how to iterate the data and blew out to 220 seconds, 462k input tokens, and 29 calls. We rejected that path and kept the transformation in Zig.

Pre-slim

28.0s · 112k · 7

Variant H on the Linear-shaped fixture.

Cold + slim

20.8s · 30k · 5

Variant N with no warm shape cache.

Warm + slim

14.8s · 31k · 5

Variant L after the local shape memory exists.

Live Grok 4.6 SuperGrok OAuth, one repetition, ReleaseSafe, --no-lean, synthetic Linear-shaped fixture. L and H are not a clean warm-versus-warm A/B.

Codegraff mice connect a model workbench to three tool stations and filter bulky results into compact reusable cards.
MCP tools sit around the model as explicit inputs and outputs; Graff filters bulky results into compact reusable shapes.

The evidence: useful signals, narrow claims

Bind reuse

135k → 25.3k

99.6s → 39.2s; 11 → 5 calls. One stable-reuse run.

Full RLM suite

203k → 105k

5/5 → 5/5; 161s → 136s; 26 → 22 calls.

Core suite

12/12 → 12/12

86.6s → 86.4s. Effectively no wall-time change.

Reported repetitions

1

Good for finding a direction, insufficient for a universal ranking.

Method boundary

  • Model and route: Grok 4.6 through SuperGrok OAuth.
  • Billing: flat-rate subscription in these runs; input cuts become spend cuts on a metered API route.
  • RLM suite: one repetition on five project-defined tasks.
  • MCP fixture: one repetition, ReleaseSafe, --no-lean.
  • Not disclosed for a universal ranking: repeated seeded trials, hardware, complete prompts and outputs, cache-state controls, TTFT, p50, and p95.

The thread also compared Graff with grok-build on six DeepSWE-shaped tasks at parallelism six. Both reported 5/6, and both failed label-sort. Graff reported 198s, 106k input, 24 calls, and 8.5 MB peak RSS. grok-build reported 506s, 194k input, 39 calls, and 177 MB peak RSS.

That internal run favored Graff on four deployment signals while preserving the same reported pass rate. It is not a normalized inference leaderboard. Harness prompts, tool policies, cache behavior, and process architecture differ by design.

DGM, not mysterious model training

We wrote “GDM-style” in the announcement thread. The correct acronym is DGM, for Darwin Gödel Machine (opens in a new tab). The project keeps an archive of candidate agent strategies, evaluates variants, and can promote one that passes the scoring contract.

That does not retrain Grok 4.6 or modify model weights. MCP learned slim is narrower still: it stores structural information about a tool result so the next call can carry less data. “Self-improving” only means something when a change beats a frozen baseline on protected checks and survives a rerun.

Try the current Graff release

The latest installer selects the current signed release for macOS, Linux, or Windows. Use --old to restore the structured-only loop, --rlm to force RLM visible, and GRAFF_XAI_X_SEARCH=0 to disable hosted X search.

curl -fsSL https://github.com/justrach/codegraff/releases/latest/download/install.sh | bash
graff --version
graff --model grok-4.6

Sources and reproducibility

  1. Announcement thread, August 26, 2026 (opens in a new tab)
  2. Graff v0.0.276 release notes and evaluation tables (opens in a new tab)
  3. Graff v0.0.277 release and downloads (opens in a new tab)
  4. ADR 0029: MCP inside RLM and value-free return shapes (opens in a new tab)
  5. ADR 0030: hide RLM on small turns (opens in a new tab)
  6. ADR 0031: xAI Responses hosts X search (opens in a new tab)
  7. Recursive Language Models (opens in a new tab)
  8. Speculative Programmatic Tool Calling (opens in a new tab)
  9. xAI X Search documentation (opens in a new tab)
  10. xAI prompt-cache guidance (opens in a new tab)
  11. Darwin Gödel Machine (opens in a new tab)