We ran a quality-diversity eval across 8 code-search tasks, 4 backends, 2 corpora, with Sonnet 4.6 agents doing real work. Not “pick the metric that makes us look good.” All three axes at once: quality, tokens, wall time.
codedb came out Pareto-optimal. Highest quality. Lowest wall time. Best tokens-per-quality-point. It wins or ties on quality in every behavioral niche. The two compression-based backends, lean-ctx and codedb's own LEAN mode, are both dominated.
This post is the data. Four releases since v0.2.579: a task-shaped context composer, auto-prepended codebase maps, MCP stability fixes, and the search + snapshot I/O pass that ties it together. Every chart is reproducible from the repo.
4.65
quality / 5
best across all backends
5/5
niches won
MAP-Elites quality leader
25.2s
avg wall time
fastest agent completion
3,892
tokens / point
best efficiency ratio
TL;DR
- Pareto-optimal. Quality 4.65/5, 18K tokens, 25s wall. No other backend beats codedb on all three axes.
- 5/5 quality niches. MAP-Elites grid across symbol-lookup, trace, pattern-find, comparison, and uncategorized. codedb leads every one.
- codedb_context. One tool call replaces 3 to 5. On SWE-bench lite: 4/4 recall, 12x fewer calls than raw CLI, 20 to 30x faster.
- reader.md. Auto-prepended codebase maps cut tool calls by 57% and tokens by 49% on narrow-symbol tasks.
- MCP stability. SIGPIPE handling, stale snapshot fix, .env bypass filter, BM25 NaN safety.
The evaluation
Most code-search benchmarks measure the wrong thing. Latency alone doesn't tell you whether an agent finishes a task. Token count alone doesn't tell you whether the answer is correct. We wanted all three: quality, tokens, and wall time, measured on an agent doing real code exploration.
The eval is a quality-diversity matrix (MAP-Elites style). Eight tasks across two corpora: facebook/react (6,620 files) and rust-lang/regex (285 files). Four backends: codedb, fts5_trigram, lean-ctx, and codedb's own LEAN mode. One Sonnet 4.6 agent per (task, backend) pair. Quality scored on a 5-point rubric: file correct, function correct, snippet faithful, explanation accurate, completeness.
No cherry-picking. No metric shopping. A backend is Pareto-dominant only if no other backend beats it on all three axes simultaneously.
quality × efficiency — sonnet 4.6 agent eval (8 tasks, 2 corpora)
Pareto frontier
Two backends survive the Pareto filter: codedb and fts5_trigram. Both lean-ctx and codedb_LEAN are dominated. For each, there is always another backend with higher quality, fewer tokens, and lower wall time.
codedb leads on quality (4.65 vs 4.38) and wall time (25.2s vs 36.9s). fts5_trigram leads on raw token count (17,172 vs 18,083). On tokens per quality point, the efficiency metric that accounts for answer correctness, codedb wins: 3,892 vs 3,925.
pareto frontier — quality × tokens × wall
| backend | quality | tokens | wall (s) | status |
|---|---|---|---|---|
| codedb | 4.65 | 18,083 | 25.2 | PARETO-OPTIMAL |
| fts5_trigram | 4.38 | 17,172 | 36.9 | PARETO-OPTIMAL |
| codedb_LEAN | 4.33 | 24,474 | 108.0 | dominated |
| lean-ctx | 4.25 | 21,452 | 67.8 | dominated |
The compression trap is real. lean-ctx markets “99% token savings,” but on real agentic tasks it used essentially the same total tokens as codedb (78,604 vs 78,424, a 0.2% difference). The compression is lossy enough that agents make extra probes to reconstruct context. codedb's own LEAN mode proved the same thing: 32% more tokens than the default, not fewer.
Every niche
The MAP-Elites grid groups tasks by behavioral niche: symbol-lookup, trace, pattern-find, comparison, and uncategorized. codedb wins or ties on quality in all five. fts5_trigram wins on tokens in three, but never on quality. lean-ctx wins zero niches on either axis.
MAP-Elites grid — quality by niche (best bolded)
symbol-lookuptracepattern-findcomparisonuncategorizedSearch latency
Underneath the agent eval is raw search performance. In a head-to-head shootout on the React corpus (6,620 files), codedb is faster on every query in the set, MCP-vs-MCP. Against lean-ctx, speedups range from 2.8x to 1,400x. Against codegraph, codedb is typically 3 to 8x faster on identifier queries. codegraph is competitive on raw latency (1 to 8ms), but its cold index costs 15.12s vs codedb's 1.18s. That's 12.8x slower, at 2.8x more disk.
warm query latency — React corpus (p50, ms)
forwardRef153x vs leanFiber160x vs leanLane250x vs leanflushPassiveEffects348x vs leanReactDOMRoot1,056x vs leanMCP stdio, 25 iterations, macOS arm64. React corpus (6,620 files).
cold index build — React corpus (6,620 files, 26.5 MB)
codedb indexes 12.8x faster than codegraph, at 2.8x less disk.
At the engine level (no MCP, no formatting), codedb's word-index lookup is 5 to 200x faster than SQLite FTS5 trigram. A forwardRef lookup costs ~1μs. A negative query costs effectively zero. Compared to what agents feel when shelling out to rg or grep, codedb is 500 to 50,000x faster.
codedb_context: one call replaces 3 to 5
v0.2.5815 shipped codedb_context, a task-shaped context composer. Your agent calls search, then outline, then symbol, then callers. codedb_context returns all of that in a single round-trip. Ranked files, inlined body lines for small definitions, callers section pre-resolved.
On SWE-bench lite file localization (4 instances, deterministic oracle), codedb_context hit 4/4 recall with 3/4 top-1 accuracy. It used 12x fewer calls than the raw CLI and finished 20 to 30x faster than every other backend. codegraph's own context mode only managed 2/4 recall. It missed half the target files.
codedb_context on SWE-bench lite (4/4 recall, 3/4 top-1)
vs codedb CLI
12x fewer
34x faster
2.6x fewer
recall 4/4
vs lean-ctx
4x fewer
22x faster
2.1x fewer
recall 4/4
vs fts5_trigram
6x fewer
20x faster
1.8x fewer
recall 4/4
vs codegraph_ctx
2.3x fewer
11x faster
3.5x fewer
recall 2/4
codedb_context query="how does setState trigger a re-render?" → 5 files ranked, 3 symbol definitions inlined src/react-reconciler/ReactFiberHooks.js:1847 dispatchSetState src/react-reconciler/ReactFiberWorkLoop.js:422 scheduleUpdateOnFiber src/react-reconciler/ReactFiberClassComponent.js:198 enqueueSetState ── callers ── ReactFiberHooks.js:1891 → dispatchSetState (via mountState) ReactFiberHooks.js:1903 → dispatchSetState (via updateState)
reader.md: codebase maps for agents
v0.2.5817 introduced reader.md, a hash-stable, auto-prepended codebase map. Think of it as an index card for the repo: the top symbols, the hot files, the dependency shape. Generated once per snapshot and prepended to codedb_context responses. Your agent gets structural awareness before it makes its first tool call.
We ran a controlled A/B against main with Sonnet 4.6 agents. On narrow-symbol tasks (the kind agents do most often), the branch with reader.md used 57% fewer tool calls and 49% fewer tokens. On broader exploration tasks, it broke even. It never made things worse. The one-time generation cost (~31K tokens) pays for itself after ~3 tasks.
task calls wall tokens T2 regex -77% -70% -54% ← narrow-symbol, biggest win T3 react -46% -21% +4% ← broad exploration, neutral T1 flask 0% 0% +11% ← small corpus, no change ───────────────────────────────────── average -41% -30% -13%
MCP stability
v0.2.5818 closes four correctness issues that showed up under real agent workloads.
SIGPIPE + broken stdout. When an MCP client disconnects mid-response, the server now ignores SIGPIPE and detects the broken pipe instead of crashing. No more dead processes after an agent closes a session mid-tool-call.
Issue-44: stale snapshots. After working tree changes, snapshot content could become invisible to search. The snapshot path now invalidates correctly on file-system events.
.env bypass. The sensitive-path filter now blocks all .env variants: .env-local, .env_production, and similar. No more accidental exposure through creative naming.
BM25 NaN safety. Zero-length documents no longer produce NaN scores that could corrupt result rankings.
Performance
The cumulative speedups since v0.2.572 stack up. Identifier token splitting in the word index gives a 3.9x geo-mean on sub-token queries. The parallel trigram build from the Zig 0.16 migration gives ~27% on cold runs. codedb_status is 9.4x faster after caching approxIndexSizeBytes with a 5s TTL.
cumulative performance gains (v0.2.572 → v0.2.5818)
sub-token search (geo-mean)
0.60ms → 0.16ms
cold find TrigramIndex
51.3ms → 40.0ms
cold outline explore.zig
51.8ms → 40.8ms
codedb_status
9.4ms → 1.0ms
v0.2.5818 added per-tier search breakdown telemetry (OTEL-style spans) and a search + snapshot I/O optimization pass. The trigram-index file-size cap was lifted from 64KB to 1MB. Wider recall for large files, no measurable latency regression.
releases since last post
v0.2.5815May 21- codedb_context composer (1 call replaces 3–5)
- Trigram cap 64KB → 1MB
- codedb_status 9.4x faster
v0.2.5816May 22- codedb read CLI subcommand
- Tier 5 full-scan short-circuit
v0.2.5817May 23- reader.md auto-prepended codebase maps
- Path traversal + DoS fixes
- Callers section in context
v0.2.5818May 25- MCP SIGPIPE + broken stdout fix
- Issue-44 stale snapshot fix
- .env bypass filter
- BM25 NaN safety
Upgrade to 0.2.5818
This is the most thoroughly evaluated release we've shipped. Pareto-optimal retrieval, a context composer that replaces multi-step workflows with a single call, and auto-generated codebase maps that cut agent token budgets in half. Four stability fixes make it safe for long-running MCP sessions.
Install or update
macOS binaries are codesigned and notarized. Linux binaries are statically linked against musl, so there is nothing to install alongside. The installer auto-registers codedb as an MCP server in Claude Code, Codex, Cursor, and Gemini CLI. For Windsurf, Devin, and Kilo Code, point your MCP config at codedb mcp.
curl -fsSL https://codedb.codegraff.com/install.sh | sh # or if you already have it: codedb update
Full benchmark data and reproduction scripts are in benchmarks/search-shootout.