0.61
Cross-file Analysis F1
0.28 → 0.61496
Real Agent Traces
Observation Kernel~90×
Faster Incremental Updates
10.78s → 0.12s53.8%
Hard-case Pass Rate
+7.6ppMy role and contributions
On a 3-person team I owned tool governance and optimization: I analyzed 793 schema-validation failures across 496 real agent traces, designed tool contracts, evidence slots and AST hints, and built trace capture and replay.
Background
RepoMind is a multilingual code knowledge base and MCP tool system from a Tencent × SZTU joint project.
Challenge
Agents routinely bypassed MCP tools to read files directly. Even when they found candidate code, they often lacked evidence explaining its relevance.
We collapsed the failures into three root causes: weak repository understanding, weak context retrieval, and no observability.
Solution
I designed a three-stage governance roadmap that reshaped retrieval from "direct search" into "candidate recall → deep read → relationship expansion → evidence verification".
Candidate recall
Alias normalization and a candidate ledger replace a broad schema rewrite — stop the bleeding first.
Deep read and expansion
Cursor pagination + evidence slots turn 'found the code' into 'explained why it matters'.
Evidence verification
A gateway, coverage checks, and a policy layer guard the critical path and zero out POL/DUP signals.
System path
Tree-sitter parses the repository and updates the code graph incrementally. Retrieval supplies code candidates to the agent through MCP tools, while the trace system records calls for diagnosis and replay.
Technical details
Incremental code graph
- Tree-sitter parsing with Merkle snapshots and freshness gating to control stale indexes.
- Single-file update time fell from 10.78s to 0.12s (~90× faster).
Agent observability
- Trajectory capture across tool calls.
- Tool-call analysis to pinpoint failure modes.
- A replay system for regression validation.
Adaptive retrieval
Key decision
Why lightweight alias normalization instead of a broad schema rewrite? It is cheap, reversible, and stops the bleeding first — reserving effort for the more critical evidence path.
- Lightweight alias normalization instead of a large-scale schema rewrite.
- Candidate ledger + cursor pagination + evidence slots.
- AST hints and RRF for better context grounding.
Results
Hard-case pass rate
D15 block recall
The pass rate is measured across 13 hard concurrent cases.
- POL/DUP signals zeroed; ~6,100 lines of governance, evidence and AST code.
Why it matters
Trace replay and evidence checks make tool-governance outcomes reproducible and measurable.
Lessons learned
Tool governance cannot replace the reasoning layer, nor can it intercept every external harness call (Read, Grep, Bash). The next stage needs a clearer boundary between governance and reasoning — and better baseline samples before expanding scope.