Wang Jianjun
All projects

RepoMind · Code Intelligence & Evaluation

Help coding agents find relevant code in large repositories and explain it with evidence.

I led trace replay and tool governance; cross-file analysis F1 rose from 0.28 to 0.61.

View diagram (swipe sideways)
RepoMind · Code Intelligence & Evaluation cover diagram
Agent Evals & Code Intelligence LeadTeam: 3Mar – Sep 2026GitHub
AgentMCPCode Graph+5

Key results

  • Raised cross-file analysis F1 from 0.28 to 0.61
  • Replayable Observation Kernel and staged evaluation across 496 traces
More results
  • Tree-sitter incremental code graph and Freshness Gate (~90× faster updates)
  • Hard-case pass rate 46.2% → 53.8%; D15 block recall 0.083 → 0.250

0.61

Cross-file Analysis F1

0.28 → 0.61

496

Real Agent Traces

Observation Kernel

~90×

Faster Incremental Updates

10.78s → 0.12s

53.8%

Hard-case Pass Rate

+7.6pp

My role and contributions

On a 3-person team I owned tool governance and optimization: I analyzed 793 schema-validation failures across 496 real agent traces, designed tool contracts, evidence slots and AST hints, and built trace capture and replay.

Background

RepoMind is a multilingual code knowledge base and MCP tool system from a Tencent × SZTU joint project.

Challenge

Agents routinely bypassed MCP tools to read files directly. Even when they found candidate code, they often lacked evidence explaining its relevance.

We collapsed the failures into three root causes: weak repository understanding, weak context retrieval, and no observability.

Solution

I designed a three-stage governance roadmap that reshaped retrieval from "direct search" into "candidate recall → deep read → relationship expansion → evidence verification".

01

Candidate recall

Alias normalization and a candidate ledger replace a broad schema rewrite — stop the bleeding first.

02

Deep read and expansion

Cursor pagination + evidence slots turn 'found the code' into 'explained why it matters'.

03

Evidence verification

A gateway, coverage checks, and a policy layer guard the critical path and zero out POL/DUP signals.

System path

Tree-sitter parses the repository and updates the code graph incrementally. Retrieval supplies code candidates to the agent through MCP tools, while the trace system records calls for diagnosis and replay.

Technical details

Incremental code graph

  • Tree-sitter parsing with Merkle snapshots and freshness gating to control stale indexes.
  • Single-file update time fell from 10.78s to 0.12s (~90× faster).

Agent observability

  • Trajectory capture across tool calls.
  • Tool-call analysis to pinpoint failure modes.
  • A replay system for regression validation.

Adaptive retrieval

Key decision

Why lightweight alias normalization instead of a broad schema rewrite? It is cheap, reversible, and stops the bleeding first — reserving effort for the more critical evidence path.

  • Lightweight alias normalization instead of a large-scale schema rewrite.
  • Candidate ledger + cursor pagination + evidence slots.
  • AST hints and RRF for better context grounding.

Results

Hard-case pass rate

46.2%53.8%

D15 block recall

0.0830.250

The pass rate is measured across 13 hard concurrent cases.

  • POL/DUP signals zeroed; ~6,100 lines of governance, evidence and AST code.

Why it matters

Trace replay and evidence checks make tool-governance outcomes reproducible and measurable.

Lessons learned

Tool governance cannot replace the reasoning layer, nor can it intercept every external harness call (Read, Grep, Bash). The next stage needs a clearer boundary between governance and reasoning — and better baseline samples before expanding scope.