Blog / Engineering

78% fewer tokens: how we made muonry edits 4x more efficient

Rach Pradhan · 6 min read

The typical AI edit workflow is 4 MCP calls: search for the code, read the function, edit it, verify the diff. We collapsed it into 1 call. Three changes to the muonry MCP daemon cut the typical edit workflow from 4 round-trips (~871 tokens) to 1 round-trip (~194 tokens). Here's what we changed and why it matters for token cost.

TL;DR

Three changes to the muonry MCP daemon cut the typical edit workflow from 4 round-trips (~871 tokens) to 1 round-trip (~194 tokens). Scope-aware edits handle search+resolve+edit+diff in a single call. Symbol caching eliminates redundant file parsing. Batch coalescing prevents line drift on same-file edits. All changes are backward-compatible.

The problem: 4 calls to edit one function

When Claude edits code through muonry, the typical flow looks like this:

Before: 4 sequential MCP calls
1. mcp__muonry__search  →  find the pattern, get file:line
2. mcp__muonry__read    →  read the function body  
3. mcp__muonry__edit    →  replace the function
4. mcp__muonry__diff    →  verify the change landed

Each call is a JSON-RPC round-trip over stdio. Each response includes MCP framing, tool metadata, and the actual content. The contentof calls 2 and 4 is almost entirely redundant — the edit handler already reads the file internally to apply the patch, and already knows the diff because it just wrote the new content.

We measured the cost: ~871 estimated tokens per edit cycle, with 28.7msof wall-clock latency across 4 calls. For a typical coding session with 20–40 edits, that's 17K–35K tokens spent just on the edit ceremony, not the actual code reasoning.

Change 1: Scope-aware edit

The biggest win. We added a pattern parameter to the edit handler. Instead of searching for code then editing, you tell the edit handler what to find and it resolves the scope internally.

After: 1 MCP call
mcp__muonry__edit: {
  "file": "src/auth.zig",
  "pattern": "validateToken",
  "content": "fn validateToken(tok: []const u8) !bool {\n    ...\n}\n"
}

What happens inside that single call:

  1. SIMD pattern search — the same matcher used by mcp__muonry__search scans the file for the pattern
  2. Symbol extraction— all functions/structs/classes are parsed from the file (up to 1024 symbols)
  3. Binary search for enclosing scope— finds the symbol whose line range contains the pattern match. Falls back to a brace-based scope resolver if no named symbol is found
  4. Patch application— replaces the entire scope with the provided content using the same atomic patcher
  5. Auto-diff — runs git diff internally, strips headers, includes the compact diff in the response
  6. Patch-and-return— the written lines are echoed back so the caller sees exactly what was written

The response includes everything the caller needs: success status, line counts, the written content, and the git diff. No follow-up calls required.

Measured results

We benchmarked both flows against a 28-line Zig file with 4 functions, editing the processData function. Each scenario ran 5 times, averaged.

Old (4 calls)
New (1 call)
Request bytes69% fewer
659 bytes
206 bytes
Response bytes80% fewer
2,828 bytes
571 bytes
Latency52% fewer
28.7 ms
13.7 ms
Est. tokens78% fewer
871 tokens
194 tokens

Benchmark: editing processData() in a 28-line Zig file. 5 runs averaged. muonry Debug build.

MetricOLD (4 calls)NEW (1 call)Change
Latency28.7ms13.7ms52% faster
Request bytes65920669% fewer
Response bytes2,82857180% fewer
Est. tokens~871~19478% fewer
Round-trips4175% fewer

Change 2: Symbol cache

When mcp__muonry__search runs with scope=true, it already parses every symbol in the file to resolve enclosing scopes. That work was being thrown away. Now it's cached.

A new 8-slot LRU symbol_cache stores up to 256 extracted symbols per file. When the edit handler needs to resolve a symbol name, it checks the cache first — O(1) lookup instead of re-reading the file and re-parsing all symbols.

Search → Edit with cache hit
# Search populates symbol cache automatically
mcp__muonry__search: {"pattern": "TODO", "path": "src/", "scope": true}

# Edit checks cache first — skips file read + parse
mcp__muonry__edit: {"file": "src/auth.zig", "symbol": "handleAuth", "content": "..."}
                        ↑ cache hit: 0µs vs 30µs cold parse

Measured improvement: 48% faster and 31% fewer bytes for the search-then-edit flow (2 calls instead of 3, with the diff folded into the edit response).

The cache is invalidated atomically whenever a file is written through muonry. No stale data, no consistency issues.

Change 3: Batch auto-coalesce

This one's a correctness fix that also improves throughput. When a batch contains multiple edits to the same file, muonry's conflict detector serializes them — they run one after another. The problem: if edit #1 changes the line count, edit #2's line numbers are wrong.

The line drift problem
batch: [
  edit lines 5-10 (remove 3 lines)  →  executes first
  edit lines 20-25                  →  WRONG: line 20 is now line 17
]

The fix: reorderSameFileEdits() runs after conflict detection. It finds same-file edit chains across serialized batches, parses each edit's range, and sorts them descending — highest line numbers first.

After: bottom-up execution
batch: [
  edit lines 20-25                  →  executes first (higher range)
  edit lines 5-10 (remove 3 lines)  →  executes second (unaffected)
]

Editing from the bottom up means each edit's line changes don't affect any subsequent edit's line numbers. This is transparent — no API change, the caller doesn't need to pre-sort. muonry handles it.

Change 4: Auto-diff in edit response

Every successful edit now includes a compact diff field in the response. The handler runs git diff --no-color -- FILE internally, strips the git headers (keeps from the first @@), and includes it alongside the patch-and-return content.

Edit response (trimmed)
{
  "ok": true,
  "op": "replace",
  "lines_removed": 8,
  "lines_added": 3,
  "content": "fn processData(d: []const u8) !void {\n    ...\n}\n",
  "diff": "@@ -9,8 +9,3 @@\n-    if (data.len == 0) ...\n+    ...\n"
}

Graceful degradation: if the file isn't git-tracked, or git fails, or it's a dry run, the diff field is simply absent. No error, no behavior change.

What changed where

FileChange+/-
handler_edit.zigpattern param, symbol cache, auto-diff+141/-11
dispatcher.zigreorderSameFileEdits()+93
symbol_cache.zig8-slot LRU symbol cache (new)+122
handler_search.zigsymbol_cache.store() after scope resolution+5

The invalidation chain ensures consistency: when an edit writes a file, both the file data cache and the symbol cache are invalidated atomically. No stale reads, no stale symbol lookups.

The takeaway

Most of the token cost in AI coding isn't the reasoning — it's the ceremony. Search, read, edit, verify. Each step produces JSON responses with metadata, framing, and content that the model already has or doesn't need.

The fix isn't compression or smaller models. It's eliminating round-trips. If the edit handler already reads the file to apply the patch, why make the caller read it first? If it just wrote the file, why make the caller diff it?

Three changes, zero API breaks, 78% fewer tokens. The best optimization is the one that makes the work disappear.

Published March 14, 2026 — muonry v0.1.03