codedb v0.2.572: 10x faster indexing, 83% less memory

Rach Pradhan · 6 min read

You know the pause. You open a project, type a search, and wait. The cursor blinks. Somewhere, your tool is reading every file from scratch, building an index it will throw away the moment you close the window.

This release is about eliminating that pause. codedb v0.2.572 indexes 10x faster, uses 83% less memory at rest, and searches with vector instructions that process 16 bytes in a single CPU cycle. Initial indexing dropped from 3.6 seconds to 346 milliseconds. Cold RSS dropped 83%. Warm RSS dropped 92%.

TL;DR

18 performance improvements, 11 bug fixes, and the biggest memory reduction in codedb's history. On openclaw (6,315 files), initial index drops from 3.6s to 346ms, cold RSS from ~3.5GB to ~580MB, warm RSS from ~1.9GB to ~150MB. The search engine now uses SIMD scanning and tiered lookup. Git subprocess spam is down 87%. MCP zombie sessions are gone.

The numbers

Before anything else, here is what shipped. These are from the release binaries on the same corpus, same machine.

Metric0.2.560.2.572Delta
Initial index time3.6 s346 ms10x faster
Cold RSS~3.5 GB~580 MB−83%
Warm RSS~1.9 GB~150 MB−92%
Git subprocesses / 30s152−87%
Trigram search55 ms53 ms−4%
Word index35 ms32 ms−9%

Benchmark: openclaw/openclaw, 6,315 files, Apple M4 Pro, ReleaseFast.

10x faster initial indexing

When multiple CPU cores index files in parallel, they have to combine their results somewhere. In the old architecture, that somewhere was a single shared data structure that every worker fought over. The more cores you threw at it, the more time they spent waiting for each other.

Before — shared bottleneck
Worker 1
Worker 2
Worker 3
Shared merge (contention)
After — worker-local
Worker 1
Worker 2
Worker 3
Partial idx
Partial idx
Partial idx
Deterministic merge (no contention)

In 0.2.572, each worker maintains its own index fragment and builds it independently. At the end, a single deterministic merge combines them with no lock contention. The result is reproducible across runs. On top of that: single-pass scan+trigram eliminates file re-reads, pre-sized trigram HashMaps save 60ms of rehashing, and the trigram extraction itself now runs in parallel.

Initial index time
0.2.56
3,600 ms
0.2.572
346 ms

The memory diet

Every search hit in the index is stored as a WordHit struct. The old struct wasted space: 8 bytes for a line number that never exceeds 2 billion, and 12 bytes of alignment padding that existed only because the compiler demanded it.

Before24 bytes per hit
file_id: u32
line_number: u64
padding
↓
After8 bytes per hit-67%
file_id: u32
line: u31 + flag: u1
saved

The old struct wasted 12 bytes on alignment padding. A packed struct with u31 line numbers fits everything in 8 bytes, cutting warm RSS by 92%.

A packed struct with u31 line numbers fits everything in 8 bytes. That single change accounts for most of the 92% warm RSS reduction. Every MCP session, every secondary project load, every short-lived CLI call was carrying that overhead.

Cold RSS improvements come from a different set of changes: staggered word index and trigram builds using c_allocator and page_allocator, worker-local arenas that are eagerly freed after indexing, and skipping the content cache for repos beyond 1,000 files.

0.2.56
0.2.572
Cold RSS
0.2.56
3,500 MB
0.2.572
580 MB
Warm RSS
0.2.56
1,900 MB
0.2.572
150 MB

RSS values from macOS /usr/bin/time -l. WordHit compaction (24B to 8B packed struct) is the primary driver of the warm reduction. Cold savings come from staggered builds, eager arena frees, and c_allocator for trigram construction.

Tiered lookup

Why scan every file in a project when most answers are sitting in an index? The tiered lookup system starts with the cheapest possible check and only escalates when it has to. Think of it as a funnel: most queries resolve at the wide top and never reach the expensive bottom.

Exact word matches resolve at Tier 0 via the word index without touching file content at all. Trigram covering sets handle most multi-character patterns at Tier 1. Only when those miss does the engine fall through to SIMD content scanning and beyond. Candidates are sorted smallest-first so tight budgets hit early exits. The searched HashMap is deferred and only allocated when Tiers 2–5 are reached, which for most identifier lookups is never.

← Most queries exit here
T0Word index62%<1 ms
T1Trigram set28%~2 ms
T2SIMD scan7%~10 ms
T3Sparse fallback2%~20 ms
T4Case-insensitive0.8%~35 ms
T5Full fallback0.2%~55 ms

Each tier gates the next. 90% of queries resolve at T0 or T1 without ever touching file content. The searched HashMap is only allocated when Tiers 2-5 are reached.

Quieter watcher

The file watcher used to fork git rev-parse on every check. That is roughly 15 subprocesses every 30 seconds on an idle repo. Now it stats .git/HEAD mtime first and only forks when something actually changed. Steady-state drops to about 2 subprocesses per 30 seconds.

MCP sessions also got a cleanup. A 10-minute idle timeout plus POLLHUP detection on stdin means zombie sessions are reaped, and dead clients trigger immediate clean shutdown instead of lingering.

Git subprocesses per 30 seconds (idle repo)
0.2.56
15
0.2.572
2

Correctness fixes

The trigram index had an unbounded growth bug: id_to_path grew with every file ever indexed, even after deletion. Now it reuses freed doc_id slots from a free-list and only grows to the peak live file count.

Other fixes: ghost entries in removeFile where stale path_to_id entries remained, an O(n) linear scan in PostingList.removeDocId replaced with O(log n) binary search, an ArrayList leak on the mmap overlay error path, and an errdefer gap that could leave word and trigram indexes out of sync on OOM. Snapshot loading now treats corrupt OUTLINE_STATE as empty rather than fatal, and config rewrites use atomic write-sync-rename.

Current recommended releasev0.2.572

Upgrade to 0.2.572

If you are on anything older than 0.2.57, the indexing and memory improvements alone are reason enough to move. The SIMD search patches and tiered lookup in 0.2.572 are the hotfixes that make the search path match the rest of the release.

Install or update

macOS binary is signed with Developer ID and notarized by Apple. Linux x86_64 binary is also available. Both are Gatekeeper-ready.

Install / update
curl -fsSL https://codedb.codegraff.com/install.sh | bash

# or if you already have it:
codedb update