codedb · retrieval experiments

We tested four ways to improve CodeDB retrieval. One earned its way in.

Rach Pradhan · 9 min read

A search result can look smarter and still put the right file farther away. That is what happened when we tested role-specific path preferences, reciprocal-rank fusion, and graph weights against CodeDB's plain lexical ranking.

The unglamorous baseline won. On 48 held-out chronological queries, every stateful or fused ranking proposal made the early result order worse. A different graph did help: following one explicit Markdown link improved Recall@10 from 0.20 to 0.30 with effectively flat warm latency.

So we shipped the bounded documentation graph and typed provenance. We did not ship the ranking ideas. This is the record of that decision, including what CodeDB users should turn on and what they should leave alone.

48

held-out queries

chronological, cross-corpus

3

ranking ideas rejected

all lost early precision

+10 pt

Recall@10

one document hop

+0.007 ms

warm latency

0.494 → 0.501 ms

The evaluation: yesterday predicts tomorrow

Random train/test splits let future vocabulary leak backward. We instead built 156 Conventional Commit queries in chronological order and reserved the latest 48 for held-out evaluation. Each query had to recover the files changed by that commit from the repository state available before it.

The ranking corpus covered CodeDB, Gin, and fd. A fourth repository, itsdangerous, produced no usable Conventional Commit queries, so it stayed in the broader parser and graph probes but was excluded from ranking macro averages. That exclusion was decided by data availability, not by a method's score.

Primary gate

MRR asks how early the first correct file appears. nDCG@10 rewards ranking all relevant files near the top. Recall@5 and Recall@10 measure whether the right files appear at all. A method had to improve the useful top of the list, not merely place one more answer near position ten.

Three plausible ideas did not improve retrieval

Held-out ranking quality

MRR on 48 chronological queries; higher is better

baseline = 100%
Lexical baseline
keep0.617

nDCG@10 0.596 · Recall@5 0.662

Role path prior
reject0.590

nDCG@10 0.573 · Recall@5 0.638

Weighted RRF
reject0.540

nDCG@10 0.536 · Recall@5 0.509

Static graph weight
neutral0.617

nDCG@10 0.596 · Recall@5 0.662

Online adaptive graph
reject0.551

nDCG@10 0.534 · Recall@5 0.551

  • Role-scoped path profiles lowered MRR from 0.617 to 0.590 and nDCG@10 from 0.596 to 0.573. The prior was not strong enough to justify hidden per-role retrieval state or another configuration layer.
  • Weighted reciprocal-rank fusion combined multiple ranked lists, but MRR fell to 0.540 and Recall@5 fell from 0.662 to 0.509. Recall@10 rose slightly to 0.711, which meant it found a few more files late while pushing the most useful answers down.
  • Online adaptive graph weighting lowered MRR to 0.551 and nDCG@10 to 0.534. Static graph weighting exactly matched the baseline, which is not a win when it adds machinery.

Why reranking could not rescue a missing candidate

Imagine sorting a shortlist of books after the librarian left the right book in the stacks. A better sorter cannot recover it. That is the candidate-recall problem: graph weights and fusion can reorder the candidates they receive, but they cannot rank a file the first-stage search never retrieved.

The neutral static result made this visible. The graph had signal, but applying it after lexical candidate generation left the aggregate metrics unchanged. The adaptive version then added noisy feedback on top and made the order worse. The next ranking experiment should improve candidate generation first, then test reranking instead of hiding another prior in the scoring path.

What earned its way in: explicit document traversal

Documentation is different from source ranking. A Markdown link is an author-written relationship: this design note points to that benchmark, or this guide continues in that reference. CodeDB now indexes those edges in a graph separate from language imports and follows them only when the caller asks.

One linked-document hop

10 documentation queries, warm measured runs

MRR

0 hops
0.133
1 hop
0.148

Recall@10

0 hops
0.20
1 hop
0.30

Warm latency

0 hops
0.494 ms
1 hop
0.501 ms

The test is deliberately small: 10 linked-document queries, with one warm-up followed by one measured run. One hop improved MRR from 0.133 to 0.148 and Recall@10 by 10 percentage points. Warm mean latency moved from 0.494 ms to 0.501 ms.

A cross-corpus parser probe found 837 links across 88 Markdown nodes. Gin alone had 561 links in eight Markdown files, so traversal is capped at two hops and 64 live files; parsing is capped at 1,024 links per file. External URLs, images, traversal paths, code spans, and malformed targets are rejected conservatively.

On a duplicated 2,868-file indexing corpus, the implementation changed wall time by −0.32% and peak RSS by +0.98%. Both stayed well inside CodeDB's 10% regression gate.

Typed provenance worked, but it is not free

Agents sometimes need more than prose. They need to know whether a result came from a parsed symbol, lexical scope, graph edge, ranked file, exact source range, or generated explanation. codedb_context format=json now emits those classes with project-relative paths, line ranges, source paths, validation state, stable omissions, and token-budget metadata.

All 12 cross-corpus provenance probe runs succeeded. Median parse coverage was 1.0. The trade-off was size: median modeled JSON was 1,506.5 bytes versus 376 raw bytes, a 3.09× metadata overhead. That is why compact Markdown remains the default and typed JSON is explicit.

What to use and what not to use

OptionDecisionReason
Default lexical rankingUseBest held-out early ranking quality
document_hops: 1Use for docsImproved linked-document recall; opt-in
document_hops: 2Use sparinglyOnly when the first hop is not enough
edge_type: documentsUse for graph inspectionShows Markdown links, separate from imports
format: jsonUse for automationTyped provenance; about 3.09× more metadata
Role path profilesDo not useRegressed MRR and nDCG@10
Weighted RRFDo not useRecall@5 and early ordering regressed
Adaptive graph rankingDo not useRegressed every primary ranking metric
# Documentation-heavy task
codedb_context task="trace the snapshot design" document_hops=1

# Inspect Markdown links directly
codedb_deps path="docs/snapshot.md" edge_type=documents

# Machine-readable evidence for another tool
codedb_context task="trace snapshot validation" format=json

For ordinary source tasks, do nothing. The lexical baseline remains the default. Add one document hop when the answer is likely spread across linked Markdown. Use two only when the first hop is visibly insufficient. Ask for JSON when software, not a person, needs to inspect the evidence.

The rule going forward

A retrieval feature does not ship because it is clever. It ships when held-out queries improve, the cost stays bounded, and users can understand when it applies. By that rule, explicit document traversal passed. Role profiles, weighted RRF, and adaptive graph ranking did not.

The implementation and acceptance evidence are in CodeDB issues #685 and #688. The negative experiment records are in #686, #687, and #689. The shipped code is in commit cc33a89.