The typical AI edit workflow is 4 MCP calls: search for the code, read the function, edit it, verify the diff. We collapsed it into 1 call. Three changes to the muonry MCP daemon cut the typical edit workflow from 4 round-trips (~871 tokens) to 1 round-trip (~194 tokens). Here's what we changed and why it matters for token cost.
TL;DR
Three changes to the muonry MCP daemon cut the typical edit workflow from 4 round-trips (~871 tokens) to 1 round-trip (~194 tokens). Scope-aware edits handle search+resolve+edit+diff in a single call. Symbol caching eliminates redundant file parsing. Batch coalescing prevents line drift on same-file edits. All changes are backward-compatible.
The problem: 4 calls to edit one function
When Claude edits code through muonry, the typical flow looks like this:
1. mcp__muonry__search → find the pattern, get file:line 2. mcp__muonry__read → read the function body 3. mcp__muonry__edit → replace the function 4. mcp__muonry__diff → verify the change landed
Each call is a JSON-RPC round-trip over stdio. Each response includes MCP framing, tool metadata, and the actual content. The contentof calls 2 and 4 is almost entirely redundant — the edit handler already reads the file internally to apply the patch, and already knows the diff because it just wrote the new content.
We measured the cost: ~871 estimated tokens per edit cycle, with 28.7msof wall-clock latency across 4 calls. For a typical coding session with 20–40 edits, that's 17K–35K tokens spent just on the edit ceremony, not the actual code reasoning.
Change 1: Scope-aware edit
The biggest win. We added a pattern parameter to the edit handler. Instead of searching for code then editing, you tell the edit handler what to find and it resolves the scope internally.
mcp__muonry__edit: {
"file": "src/auth.zig",
"pattern": "validateToken",
"content": "fn validateToken(tok: []const u8) !bool {\n ...\n}\n"
}What happens inside that single call:
- SIMD pattern search — the same matcher used by
mcp__muonry__searchscans the file for the pattern - Symbol extraction— all functions/structs/classes are parsed from the file (up to 1024 symbols)
- Binary search for enclosing scope— finds the symbol whose line range contains the pattern match. Falls back to a brace-based scope resolver if no named symbol is found
- Patch application— replaces the entire scope with the provided content using the same atomic patcher
- Auto-diff — runs
git diffinternally, strips headers, includes the compact diff in the response - Patch-and-return— the written lines are echoed back so the caller sees exactly what was written
The response includes everything the caller needs: success status, line counts, the written content, and the git diff. No follow-up calls required.
Measured results
We benchmarked both flows against a 28-line Zig file with 4 functions, editing the processData function. Each scenario ran 5 times, averaged.
Benchmark: editing processData() in a 28-line Zig file. 5 runs averaged. muonry Debug build.
| Metric | OLD (4 calls) | NEW (1 call) | Change |
|---|---|---|---|
| Latency | 28.7ms | 13.7ms | 52% faster |
| Request bytes | 659 | 206 | 69% fewer |
| Response bytes | 2,828 | 571 | 80% fewer |
| Est. tokens | ~871 | ~194 | 78% fewer |
| Round-trips | 4 | 1 | 75% fewer |
Change 2: Symbol cache
When mcp__muonry__search runs with scope=true, it already parses every symbol in the file to resolve enclosing scopes. That work was being thrown away. Now it's cached.
A new 8-slot LRU symbol_cache stores up to 256 extracted symbols per file. When the edit handler needs to resolve a symbol name, it checks the cache first — O(1) lookup instead of re-reading the file and re-parsing all symbols.
# Search populates symbol cache automatically
mcp__muonry__search: {"pattern": "TODO", "path": "src/", "scope": true}
# Edit checks cache first — skips file read + parse
mcp__muonry__edit: {"file": "src/auth.zig", "symbol": "handleAuth", "content": "..."}
↑ cache hit: 0µs vs 30µs cold parseMeasured improvement: 48% faster and 31% fewer bytes for the search-then-edit flow (2 calls instead of 3, with the diff folded into the edit response).
The cache is invalidated atomically whenever a file is written through muonry. No stale data, no consistency issues.
Change 3: Batch auto-coalesce
This one's a correctness fix that also improves throughput. When a batch contains multiple edits to the same file, muonry's conflict detector serializes them — they run one after another. The problem: if edit #1 changes the line count, edit #2's line numbers are wrong.
batch: [ edit lines 5-10 (remove 3 lines) → executes first edit lines 20-25 → WRONG: line 20 is now line 17 ]
The fix: reorderSameFileEdits() runs after conflict detection. It finds same-file edit chains across serialized batches, parses each edit's range, and sorts them descending — highest line numbers first.
batch: [ edit lines 20-25 → executes first (higher range) edit lines 5-10 (remove 3 lines) → executes second (unaffected) ]
Editing from the bottom up means each edit's line changes don't affect any subsequent edit's line numbers. This is transparent — no API change, the caller doesn't need to pre-sort. muonry handles it.
Change 4: Auto-diff in edit response
Every successful edit now includes a compact diff field in the response. The handler runs git diff --no-color -- FILE internally, strips the git headers (keeps from the first @@), and includes it alongside the patch-and-return content.
{
"ok": true,
"op": "replace",
"lines_removed": 8,
"lines_added": 3,
"content": "fn processData(d: []const u8) !void {\n ...\n}\n",
"diff": "@@ -9,8 +9,3 @@\n- if (data.len == 0) ...\n+ ...\n"
}Graceful degradation: if the file isn't git-tracked, or git fails, or it's a dry run, the diff field is simply absent. No error, no behavior change.
What changed where
| File | Change | +/- |
|---|---|---|
| handler_edit.zig | pattern param, symbol cache, auto-diff | +141/-11 |
| dispatcher.zig | reorderSameFileEdits() | +93 |
| symbol_cache.zig | 8-slot LRU symbol cache (new) | +122 |
| handler_search.zig | symbol_cache.store() after scope resolution | +5 |
The invalidation chain ensures consistency: when an edit writes a file, both the file data cache and the symbol cache are invalidated atomically. No stale reads, no stale symbol lookups.
The takeaway
Most of the token cost in AI coding isn't the reasoning — it's the ceremony. Search, read, edit, verify. Each step produces JSON responses with metadata, framing, and content that the model already has or doesn't need.
The fix isn't compression or smaller models. It's eliminating round-trips. If the edit handler already reads the file to apply the patch, why make the caller read it first? If it just wrote the file, why make the caller diff it?
Three changes, zero API breaks, 78% fewer tokens. The best optimization is the one that makes the work disappear.
Published March 14, 2026 — muonry v0.1.03