muonry v0.1.15: 21 audit findings, fixed

Rach Pradhan · 6 min read

muonry v0.1.15 ships seven new safety features across read, edit, and search.

Read mode flags when its output is unannotated raw bytes. Edit surfaces pattern-match counts and warns when no positioning is given. The pipe protocol recovers from oversized commands instead of dropping the connection. Directory walks cap at 64 levels deep. ThecontentFileargument is confined to the working directory. ls reports symlink sizes correctly. Binary detection samples three regions of the file instead of just the head.

On the same workloads as Claude Code's built-in tools, the daemon uses 96% fewer tokens for outline and 69% fewer for directory listing.

TL;DR

Two audit passes against the muonry daemon turned up 26 issues across security, correctness, API contracts, and stability. v0.1.15 fixes 21 of them fully. 3 are partial fixes with documented residual scope. 2 are wontfix with a written reason. On the same workloads, muonry uses 96% fewer tokens than Claude Code's built-in Read for outline, 69% fewer than Bash for directory listing, and parity with Grep for search.

The audit method

We ran a Sonnet 4.6 agent against the muonry source with one instruction: find every silent failure mode you can. The first pass turned up 19. We fixed all 19, rebuilt the daemon, and ran the same agent against the new binary. The second pass turned up 4 more, including one case where our own fix did not cover the full code path. We fixed those, then ran a third pass to verify all 21 fixes behave as advertised on the live MCP connection.

Each finding lives as a markdown card underevals/edge-cases/with the original repro, the actual change we made, and any residual risk we chose not to address. The same prompts that produced the findings are checked into the repo, so the audit is reproducible.

The bias of the audit was strongly toward silent failures, the kind where muonry returnsok:truebut the response is wrong, or the daemon quietly exits and drops every subsequent command. Loud failures (a tool that errors on bad input) cost one round trip. Silent failures cost retries, re-reads, and human attention.

Pass 1 findings
19
Round 1 fixed
17
+ 2 partial
Pass 2 findings
4
incl. 1 hole in our own fix
Round 2 fixed
4
Pass 3 verified
21
+ 3 partial, 2 wontfix
audit findings
resolved

What we fixed

File safety on edits

Editing with a negativeaftervalue used to fall through to a no-positioning fallback that replaced the entire file. The most common path to a negative value is an off-by-one in caller code. v0.1.15 returns a clear error instead. Editing with bothrangeandafterset used to silently drop one based on internal precedence. That now errors too. Reading lines with a 0 line number used to return empty content. That now returns an explicit 1-indexed error.

Daemon stability

The pipe protocol caps each command at 1 MiB per line. The pre-fix line reader returned the same null sentinel for clean end-of-input and for an oversized command, and the dispatcher used that null to break out of the loop. Result: a single oversized command silently exited the daemon and every subsequent command in the same connection was dropped. v0.1.15 differentiates the two cases, drains the rest of the oversized line so the next read starts on a clean boundary, emits a clear error, and continues processing.

Before: one bad command kills the daemon
cmd 1: { "tool":"read", ... }
cmd 2: 1.2 MiB JSON (oversize)
daemon: pipeReadLine returns null
dispatcher: orelse break
cmd 3: { ... } (never runs)
cmd 4: { ... } (never runs)
After: error, drain, continue
cmd 1: { "tool":"read", ... }
cmd 2: 1.2 MiB JSON (oversize)
daemon: error.LineTooLong
drain to next \n; emit error response
cmd 3: { ... } (runs as expected)
cmd 4: { ... } (runs as expected)

Bounded walks

Bothsearchandsearch find=truerecursively walk a directory tree. The pre-fix versions had no depth limit and no inode dedup, so a self-referential symlink, a recursive bind mount, or any node_modules with a circular dep would spin until the 15-second cancel watchdog fired. v0.1.15 caps both walkers at depth 64. We chose a depth cap over an inode hashset deliberately. The hashset would grow with the walk and could blow up RSS on a real monorepo with hundreds of thousands of files. A 64-deep cap is constant memory, fits every real codebase we have measured, and breaks all the cycle shapes we have seen in the wild.

Before: unbounded walk
project/
src/
loop -> ../project/
project/src/loop -> ../project/
project/src/project/src/loop ...
(walks until 15s watchdog kills it)
After: depth cap = 64
0 project/
1 src/
2 loop -> ../project/
3 ... 63
64 stop
constant memory, no inode hashset

Bounded reads

Reading a file withmode=fullannotates every line with a number and emits a header. The output goes through a 1 MiB scratch buffer. The annotation overhead, about nine bytes per line, was capable of pushing the annotated output past 1 MiB even when the raw file was much smaller. A 105,000 line file at 210 KB raw came back at 1,046,529 bytes, silently truncated. v0.1.15 estimates the annotated size up front and bypasses the annotated path when the estimate exceeds the buffer, emitting raw bytes through the streaming write. The response carries a newannotated:falseflag so callers can detect the format change. The same file now returns all 210,000 bytes.

Before: silent truncation at 1 MiB
210 KB raw
+ 836 KB line-number annotations
01,046,529 bytes returned (truncated, no flag)
After: estimate first, bypass to streaming write
est = data.len + idx.count * 12 + 256
>
out.len (1 MiB)
take raw bypass; emit "annotated":false in response
210,000 bytes streamed (full file)

Path containment

ThecontentFileargument on edit and create lets a caller point at a file whose contents become the edit body. The pre-fix handler accepted any path the daemon's UID could read, which means an MCP client could read /etc/passwd, ssh keys, or anything else outside the project. v0.1.15 resolves the path through realpath and refuses anything that does not sit under the daemon's working directory.

Valid JSON on every response

The image and binary-file branches of read used to write file paths into the response without escaping. A path containing a backslash, a quote, or a control byte produced invalid JSON that broke MCP clients trying to parse it. v0.1.15 routes every path through the same JSON escaper used elsewhere in the daemon, so the response is valid JSON regardless of what the path looks like.

Better signals on partial responses

Edits that match a pattern at multiple sites used to silently edit the first hit. The response now includes the match count and a saturation flag so callers can refine the pattern instead of assuming the edit landed where they expected. Edits without any positioning set now return awhole_file_replacewarning so the default behaviour is visible. Searches that combine mutually-exclusive flags (regex plus word_boundary plus case_insensitive) now error instead of silently dropping all but one. Dry-run edits used to only guard the replace op, so delete and insert_after still dispatched. Now dry-run guards every op.

Better detection on weird inputs

Binary detection used to sniff the first 512 bytes of the file. Files with a UTF-8 header, a licence comment block, or a shebang prepended to binary content passed through as text and produced garbled output. v0.1.15 samples the head, middle, and tail of any file at least 1.5 KiB long and flags as binary if any region contains a null byte. ls used to stat symlink entries, which followed the link and reported the size of the target. v0.1.15 uses lstat so the reported size is the link itself.

Token efficiency

We measured three common operations against Claude Code's built-in tools on the same files. Tokens estimated at bytes divided by four.

muonry
built-in (Read / Grep / Bash)
Outline (752-line file)96% fewer
374 tok
8,656 tok
Search literal in dirparity
209 tok
184 tok
List dir with sizes69% fewer
165 tok
532 tok

Tokens estimated at bytes / 4. Same files in both columns. Outline target was muonry/src/handler_read.zig (752 lines).

The outline win scales with file size. A 5,000 line source file is roughly 50,000 tokens to read fully but 400 to outline. For an agent exploring a codebase,outlinefirst thensymbolfor the specific function you care about is usually a hundred times cheaper than reading whole files.

Search is roughly parity with grep. The JSON envelope adds a small fixed overhead. In return you get structured matches with line numbers and, withscope=true, the enclosing function, which grep cannot produce without follow-up reads.

ls strips inode, permissions, link count, owner, group, and timestamps, all noise an agent does not use, and returns name and size pairs. The savings come entirely from removing metadata the model would have ignored.

The audit fixes contribute to this story indirectly. Every silent failure forces the agent to recover, and every recovery is another round trip. A truncatedfullread used to push agents into a Bashcatfallback, which is four to eight times more tokens for the same content. A walk that hung used to push agents into narrowing the path and retrying. v0.1.15 closes 21 of those failure modes, so the token cost stays on the original call.

What is left

Three of the 26 findings are partial fixes. Thefullmode bypass covers the common case but other read modes (outline, lines, symbol, compact, json) still cap at 1 MiB silently. Their output is structurally bounded so saturation is uncommon in practice. Mmap now caps at 256 MiB per file, but a file truncated mid-read can still produce a SIGBUS, which would need a signal handler in the daemon to fully fix. Therangeandafterconflict check rejects non-empty range plus after, but treats an empty range string as omitted, which is intentional for wrapper compatibility.

Two are wontfix with a written reason. The search system-root deny list is, by design, a deny list rather than an allow list, because an allow list would block common user workflows. The emptyrangestring treated as omitted is also wontfix for the same wrapper-compatibility reason.

The takeaway

Each of these features is a place muonry used to be quiet about something the caller needed to know. Now it tells you. The pattern-match count is in the response. The annotation bypass is in the response. The depth cap kicks in before the watchdog does. The validation error fires before the file gets touched.

None of this changes the muonry API. Existing calls return the same shape they did in v0.1.14, plus the new fields when they apply. The default behaviour on every tool got safer without any caller-side migration.

The biggest practical win is the one you do not see directly: fewer recovery round trips. A truncated read used to push agents into a Bash catfallback at four to eight times the token cost. A walk that hung used to push agents into narrowing the path and retrying. v0.1.15 keeps the work on the original call. The cheapest tool call is the one that does what it says the first time.

Upgrade

macOS arm64 is signed and notarized. Linux x86_64 is a static ELF and runs on glibc and musl distros without extra runtime deps. Runmuonry updateto pick up the new binary. If your MCP client is connected, restart it after the update so it picks up the new daemon.

Full release notes are on the changelog. Issues go to the GitHub repo.