You know the pause. You open a project, type a search, and wait. The cursor blinks. Somewhere, your tool is reading every file from scratch, building an index it will throw away the moment you close the window.
This release is about eliminating that pause. codedb v0.2.572 indexes 10x faster, uses 83% less memory at rest, and searches with vector instructions that process 16 bytes in a single CPU cycle. Initial indexing dropped from 3.6 seconds to 346 milliseconds. Cold RSS dropped 83%. Warm RSS dropped 92%.
TL;DR
18 performance improvements, 11 bug fixes, and the biggest memory reduction in codedb's history. On openclaw (6,315 files), initial index drops from 3.6s to 346ms, cold RSS from ~3.5GB to ~580MB, warm RSS from ~1.9GB to ~150MB. The search engine now uses SIMD scanning and tiered lookup. Git subprocess spam is down 87%. MCP zombie sessions are gone.
The numbers
Before anything else, here is what shipped. These are from the release binaries on the same corpus, same machine.
| Metric | 0.2.56 | 0.2.572 | Delta |
|---|---|---|---|
| Initial index time | 3.6 s | 346 ms | 10x faster |
| Cold RSS | ~3.5 GB | ~580 MB | −83% |
| Warm RSS | ~1.9 GB | ~150 MB | −92% |
| Git subprocesses / 30s | 15 | 2 | −87% |
| Trigram search | 55 ms | 53 ms | −4% |
| Word index | 35 ms | 32 ms | −9% |
Benchmark: openclaw/openclaw, 6,315 files, Apple M4 Pro, ReleaseFast.
10x faster initial indexing
When multiple CPU cores index files in parallel, they have to combine their results somewhere. In the old architecture, that somewhere was a single shared data structure that every worker fought over. The more cores you threw at it, the more time they spent waiting for each other.
In 0.2.572, each worker maintains its own index fragment and builds it independently. At the end, a single deterministic merge combines them with no lock contention. The result is reproducible across runs. On top of that: single-pass scan+trigram eliminates file re-reads, pre-sized trigram HashMaps save 60ms of rehashing, and the trigram extraction itself now runs in parallel.
The memory diet
Every search hit in the index is stored as a WordHit struct. The old struct wasted space: 8 bytes for a line number that never exceeds 2 billion, and 12 bytes of alignment padding that existed only because the compiler demanded it.
The old struct wasted 12 bytes on alignment padding. A packed struct with u31 line numbers fits everything in 8 bytes, cutting warm RSS by 92%.
A packed struct with u31 line numbers fits everything in 8 bytes. That single change accounts for most of the 92% warm RSS reduction. Every MCP session, every secondary project load, every short-lived CLI call was carrying that overhead.
Cold RSS improvements come from a different set of changes: staggered word index and trigram builds using c_allocator and page_allocator, worker-local arenas that are eagerly freed after indexing, and skipping the content cache for repos beyond 1,000 files.
RSS values from macOS /usr/bin/time -l. WordHit compaction (24B to 8B packed struct) is the primary driver of the warm reduction. Cold savings come from staggered builds, eager arena frees, and c_allocator for trigram construction.
SIMD search engine
Imagine reading a book one letter at a time. You look at the first character, check if it matches, move to the next, check again. That is how traditional text search works at the CPU level: one byte, one comparison, one step forward.
SIMD (Single Instruction, Multiple Data) changes the economics. Instead of checking one byte, the CPU loads 16 bytes into a special wide register and compares all of them in a single instruction. One cycle, 16 answers.
SIMD (Single Instruction, Multiple Data) uses vector registers to compare 16 bytes per CPU cycle. The scanner checks all positions simultaneously instead of stepping through one at a time.
The 0.2.572 search engine uses two SIMD paths. The first is a fast-reject scanner: it checks whether the first byte of your search pattern appears anywhere in a 16-byte window. If it does not, the entire window is skipped. The second is a substring scanner that simultaneously checks the first and last bytes of the pattern to narrow candidates before doing a full comparison.
The biggest structural change is that searchInContent no longer splits files into lines before searching. It scans the raw buffer directly and only computes line numbers after finding a match. For the majority of files that contain no matches, the line-splitting work is never done at all.
Tiered lookup
Why scan every file in a project when most answers are sitting in an index? The tiered lookup system starts with the cheapest possible check and only escalates when it has to. Think of it as a funnel: most queries resolve at the wide top and never reach the expensive bottom.
Exact word matches resolve at Tier 0 via the word index without touching file content at all. Trigram covering sets handle most multi-character patterns at Tier 1. Only when those miss does the engine fall through to SIMD content scanning and beyond. Candidates are sorted smallest-first so tight budgets hit early exits. The searched HashMap is deferred and only allocated when Tiers 2–5 are reached, which for most identifier lookups is never.
Each tier gates the next. 90% of queries resolve at T0 or T1 without ever touching file content. The searched HashMap is only allocated when Tiers 2-5 are reached.
Quieter watcher
The file watcher used to fork git rev-parse on every check. That is roughly 15 subprocesses every 30 seconds on an idle repo. Now it stats .git/HEAD mtime first and only forks when something actually changed. Steady-state drops to about 2 subprocesses per 30 seconds.
MCP sessions also got a cleanup. A 10-minute idle timeout plus POLLHUP detection on stdin means zombie sessions are reaped, and dead clients trigger immediate clean shutdown instead of lingering.
Correctness fixes
The trigram index had an unbounded growth bug: id_to_path grew with every file ever indexed, even after deletion. Now it reuses freed doc_id slots from a free-list and only grows to the peak live file count.
Other fixes: ghost entries in removeFile where stale path_to_id entries remained, an O(n) linear scan in PostingList.removeDocId replaced with O(log n) binary search, an ArrayList leak on the mmap overlay error path, and an errdefer gap that could leave word and trigram indexes out of sync on OOM. Snapshot loading now treats corrupt OUTLINE_STATE as empty rather than fatal, and config rewrites use atomic write-sync-rename.
Upgrade to 0.2.572
If you are on anything older than 0.2.57, the indexing and memory improvements alone are reason enough to move. The SIMD search patches and tiered lookup in 0.2.572 are the hotfixes that make the search path match the rest of the release.
Install or update
macOS binary is signed with Developer ID and notarized by Apple. Linux x86_64 binary is also available. Both are Gatekeeper-ready.
curl -fsSL https://codedb.codegraff.com/install.sh | bash # or if you already have it: codedb update