A search result can look smarter and still put the right file farther away. That is what happened when we tested role-specific path preferences, reciprocal-rank fusion, and graph weights against CodeDB's plain lexical ranking.
The unglamorous baseline won. On 48 held-out chronological queries, every stateful or fused ranking proposal made the early result order worse. A different graph did help: following one explicit Markdown link improved Recall@10 from 0.20 to 0.30 with effectively flat warm latency.
So we shipped the bounded documentation graph and typed provenance. We did not ship the ranking ideas. This is the record of that decision, including what CodeDB users should turn on and what they should leave alone.
48
held-out queries
chronological, cross-corpus
3
ranking ideas rejected
all lost early precision
+10 pt
Recall@10
one document hop
+0.007 ms
warm latency
0.494 → 0.501 ms
The evaluation: yesterday predicts tomorrow
Random train/test splits let future vocabulary leak backward. We instead built 156 Conventional Commit queries in chronological order and reserved the latest 48 for held-out evaluation. Each query had to recover the files changed by that commit from the repository state available before it.
The ranking corpus covered CodeDB, Gin, and fd. A fourth repository, itsdangerous, produced no usable Conventional Commit queries, so it stayed in the broader parser and graph probes but was excluded from ranking macro averages. That exclusion was decided by data availability, not by a method's score.
Primary gate
MRR asks how early the first correct file appears. nDCG@10 rewards ranking all relevant files near the top. Recall@5 and Recall@10 measure whether the right files appear at all. A method had to improve the useful top of the list, not merely place one more answer near position ten.
Three plausible ideas did not improve retrieval
Held-out ranking quality
MRR on 48 chronological queries; higher is better
nDCG@10 0.596 · Recall@5 0.662
nDCG@10 0.573 · Recall@5 0.638
nDCG@10 0.536 · Recall@5 0.509
nDCG@10 0.596 · Recall@5 0.662
nDCG@10 0.534 · Recall@5 0.551
- Role-scoped path profiles lowered MRR from 0.617 to 0.590 and nDCG@10 from 0.596 to 0.573. The prior was not strong enough to justify hidden per-role retrieval state or another configuration layer.
- Weighted reciprocal-rank fusion combined multiple ranked lists, but MRR fell to 0.540 and Recall@5 fell from 0.662 to 0.509. Recall@10 rose slightly to 0.711, which meant it found a few more files late while pushing the most useful answers down.
- Online adaptive graph weighting lowered MRR to 0.551 and nDCG@10 to 0.534. Static graph weighting exactly matched the baseline, which is not a win when it adds machinery.
Why reranking could not rescue a missing candidate
Imagine sorting a shortlist of books after the librarian left the right book in the stacks. A better sorter cannot recover it. That is the candidate-recall problem: graph weights and fusion can reorder the candidates they receive, but they cannot rank a file the first-stage search never retrieved.
The candidate-recall ceiling
A reranker can only reorder files that reach it
1 · lexical candidates
snapshot.zig
store.zig
index.zig
2 · graph reranker
1. index.zig
2. snapshot.zig
3. store.zig
correct file
reader_md.zig
not in candidate set
Reordering succeeded. Retrieval still failed.
The neutral static result made this visible. The graph had signal, but applying it after lexical candidate generation left the aggregate metrics unchanged. The adaptive version then added noisy feedback on top and made the order worse. The next ranking experiment should improve candidate generation first, then test reranking instead of hiding another prior in the scoring path.
What earned its way in: explicit document traversal
Documentation is different from source ranking. A Markdown link is an author-written relationship: this design note points to that benchmark, or this guide continues in that reference. CodeDB now indexes those edges in a graph separate from language imports and follows them only when the caller asks.
One linked-document hop
10 documentation queries, warm measured runs
MRR
Recall@10
Warm latency
The test is deliberately small: 10 linked-document queries, with one warm-up followed by one measured run. One hop improved MRR from 0.133 to 0.148 and Recall@10 by 10 percentage points. Warm mean latency moved from 0.494 ms to 0.501 ms.
A cross-corpus parser probe found 837 links across 88 Markdown nodes. Gin alone had 561 links in eight Markdown files, so traversal is capped at two hops and 64 live files; parsing is capped at 1,024 links per file. External URLs, images, traversal paths, code spans, and malformed targets are rejected conservatively.
On a duplicated 2,868-file indexing corpus, the implementation changed wall time by −0.32% and peak RSS by +0.98%. Both stayed well inside CodeDB's 10% regression gate.
Typed provenance worked, but it is not free
Agents sometimes need more than prose. They need to know whether a result came from a parsed symbol, lexical scope, graph edge, ranked file, exact source range, or generated explanation. codedb_context format=json now emits those classes with project-relative paths, line ranges, source paths, validation state, stable omissions, and token-budget metadata.
One context, two output contracts
Median cross-corpus probe size
default · compact · human-readable
opt-in · paths · ranges · evidence classes
All 12 cross-corpus provenance probe runs succeeded. Median parse coverage was 1.0. The trade-off was size: median modeled JSON was 1,506.5 bytes versus 376 raw bytes, a 3.09× metadata overhead. That is why compact Markdown remains the default and typed JSON is explicit.
What to use and what not to use
| Option | Decision | Reason |
|---|---|---|
| Default lexical ranking | Use | Best held-out early ranking quality |
| document_hops: 1 | Use for docs | Improved linked-document recall; opt-in |
| document_hops: 2 | Use sparingly | Only when the first hop is not enough |
| edge_type: documents | Use for graph inspection | Shows Markdown links, separate from imports |
| format: json | Use for automation | Typed provenance; about 3.09× more metadata |
| Role path profiles | Do not use | Regressed MRR and nDCG@10 |
| Weighted RRF | Do not use | Recall@5 and early ordering regressed |
| Adaptive graph ranking | Do not use | Regressed every primary ranking metric |
# Documentation-heavy task codedb_context task="trace the snapshot design" document_hops=1 # Inspect Markdown links directly codedb_deps path="docs/snapshot.md" edge_type=documents # Machine-readable evidence for another tool codedb_context task="trace snapshot validation" format=json
For ordinary source tasks, do nothing. The lexical baseline remains the default. Add one document hop when the answer is likely spread across linked Markdown. Use two only when the first hop is visibly insufficient. Ask for JSON when software, not a person, needs to inspect the evidence.
The rule going forward
A retrieval feature does not ship because it is clever. It ships when held-out queries improve, the cost stays bounded, and users can understand when it applies. By that rule, explicit document traversal passed. Role profiles, weighted RRF, and adaptive graph ranking did not.
The implementation and acceptance evidence are in CodeDB issues #685 and #688. The negative experiment records are in #686, #687, and #689. The shipped code is in commit cc33a89.