Every file an AI agent reads could be lying to it. Bidi overrides make code appear to do one thing while actually doing another. Zero-width characters hide content from human reviewers. ANSI escapes hijack terminal output.
We wanted every read to tell the agent whether it could trust what it was seeing. So we added a review warning scanner that runs on every file, every time. Then we made the whole thing 20% faster than v0.1.12 was without the scanner.
TL;DR
v0.1.13 scans every file for suspicious content (bidi overrides, invisible unicode, ANSI escapes, control chars) and surfaces structured warnings. Despite this new work, full reads are 20% faster than v0.1.12 on medium-to-large files. Two optimizations made it possible: a SIMD fast path that skips clean files at 16 bytes per cycle, and a bulk-memcpy JSON escape that cuts buffer operations by 10x. Plus a new bigram-indexed search engine and headless MCP support.
The tradeoff we didn't want to make
The naive implementation scanned every byte of every file on every read. On our test file (1,194 lines of Zig), this added ~200 microseconds per read — more than doubling the latency for small operations like outline or symbol extraction. That's not acceptable for a tool that handles hundreds of reads per session.
So the question became: can we scan every file without paying for it? The answer turned out to be two separate optimizations, each targeting a different bottleneck.
SIMD scanner: skip clean files at 16 bytes per cycle
99% of source files are clean ASCII — no control characters, no high bytes, no surprises. Instead of checking every byte individually, the scanner loads 16 bytes at a time into a SIMD vector register and tests all of them in a single instruction. If a byte is below 0x20 (except tab and newline), equals 0x7F, or has the high bit set, the file gets flagged for detailed analysis. If the entire file passes, the expensive per-byte scan is skipped entirely.
If all 16 bytes pass (no control chars, no high bits), the entire chunk is skipped. One bad byte flags the file for a full per-byte analysis. Most source files pass every chunk and never enter the slow path.
The second optimization: warning caching. The file cache already stores raw bytes and a pre-built line index. We extended it to also cache the warning set. On cache hit, warnings come back instantly — no analysis at all. Between the SIMD fast path and the cache, the warning scanner is effectively free for the vast majority of reads.
Bulk-memcpy JSON escape
While profiling, we found the real bottleneck wasn't warning analysis — it was JSON serialization. The old pipe output function escaped content byte-by-byte, copying one character at a time into an 8KB staging buffer. For a 50KB response, that's 50,000 individual buffer writes.
Safe runs are copied in one memcpy. Only special characters (quotes, backslashes, control chars) trigger individual escape handling. For typical source code, this cuts buffer operations by ~10x.
The fix: scan forward for runs of safe characters (anything except ", \, and control chars), then memcpy the entire run in one shot. Most source code is long runs of printable ASCII interrupted by occasional newlines and quotes. This cuts buffer operations by ~10x.
Benchmarks: own repo
All measurements: p50 latency in microseconds, pipe mode, 100+ runs, warm file cache, Apple M-series. v0.1.13 includes review warning analysis on every read — v0.1.12 did not.
The top 4 operations are faster than v0.1.12 despite the new warning scanner. The bulk-memcpy escape dominates on larger responses. Outline on small files shows the remaining overhead from warning analysis (~60us), partially offset by the SIMD fast path.
External repo: opencode (4,528 files)
We benchmarked against sst/opencode — a 139MB TypeScript monorepo with 4,528 files and 1,479 .tsx components. This is closer to what muonry sees in production.
Full reads on large files are 20–23% fasteracross the board. The improvement scales with response size — the bigger the file, the more the bulk-memcpy escape saves.
New search engine: ziggram
Searching a codebase usually means scanning every file. That works, but it scales linearly — 10,000 files means 10,000 scans. v0.1.13 introduces ziggram, a reusable search engine that uses a different strategy: index every file by its 2-byte sequences (bigrams), and use the index to eliminate candidates before any content is read.
The bigram index eliminates 93% of files before any content is read. Only files containing every 2-byte sequence from the query survive to the scan phase.
When you search for handleAuth, the engine checks which files contain the bigrams ha, an, nd, etc. Files missing any bigram are eliminated instantly. In practice, this filters out 90%+ of files before any content scan happens.
On top of that: background refresh (the index stays warm without blocking queries), frecency ranking (files you access often appear first), and meta_search — a planner that generates multiple search variants (literal, case-insensitive, word-boundary, regex) and merges the results.
Everything else in v0.1.13
- Headless MCP— muonry works when spawned without a TTY. Docker, forge, CI — no more SIGKILL on startup.
- zigedit CLI— standalone high-level edit tool. Edit by symbol name, by pattern match, or by line range. 291KB binary.
- Edit/patch enforcement —
patchrejects symbol and pattern at runtime. Clear tool roles. - no_ignore flag— search inside node_modules, .git, dist when you explicitly need to.
- content_b64— base64-encoded content for shell-safe edits over SSH and Docker.
- 8MB stack savings— consolidated 9 separate 1MB buffers into one shared buffer per read request.
- 19 issues closed in a single session.
What's next
The outline regression on small files (~60us overhead) is the remaining target. The SIMD fast path eliminates the warning scan cost, but the outline generation itself got slower from the larger binary footprint. We're looking at symbol extraction caching and outline pre-computation in the file cache.
The search regression (~50% slower than v0.1.12) is from the new dependencies (ziggram + codedb + sqlite). The faster_search path with the bigram index is actually faster on large repos — the regression is only on the raw search baseline. Separate optimization target.
cd muonry && zig build -Doptimize=ReleaseFast cp zig-out/bin/muonry ~/bin/muonry codesign -s - -f ~/bin/muonry && xattr -cr ~/bin/muonry