Wang Jianjun
All projects

Agent Engineering Delivery & Evaluation Methodology

Make design-agent delivery reproducible, reviewable and diagnosable.

I worked on M0 Design Agent evaluations, Figma structured editing and MCP integration, including a benchmark of 84 tasks, 160 pages and 979 criteria.

View diagram (swipe sideways)
BUILD → VERIFY → IMPROVEAgent Delivery & EvalsFrozen inputsPRD · target · constraintsNative outputEditable artifact + evidenceIndependent reviewValidity · coverage · qualityFindings → Repair → RegressionARCHITECTURE ILLUSTRATIONWJ / SELECTED WORK
Agent Developer · Agent Systems & EvalsAug 2026 – Present
Agent SystemsEvaluation BenchmarkFigma AST / JSX+3

Key results

  • Constructed 84-task / 160-page / 979-atomic-criteria real-world business benchmark
  • Reached 95%+ core acceptance auto-coverage; dropped single-case eval time from 20-30m to 3-5m
More results
  • Figma SceneGraph ↔ JSX declarative editing protocol with Reconcile, 95%+ complex edit success rate (-60% tool calls)
  • MCP packaging and concurrency stability, raising complex generation E2E success to 95%+ and halving exceptions

95%+

Core Acceptance Coverage

20-30m → 3-5m

979

Atomic Business Criteria

84 tasks / 160 pages

95%+

Complex Edit Success Rate

-60% tool calls

95%+

MCP Generation E2E Rate

-50% concurrency errors

Background

A design agent does not produce a single text answer. It must turn a PRD and canvas context into design content that another person can continue to edit. A successful tool call, a screenshot or a nonempty node is not proof of completed delivery.

Real delivery answers three different questions: is the target correct, is the artifact usable with complete evidence, and were requirements satisfied? Collapsing them into one success state hides failures and provides no useful iteration path.

Solution

01

Freeze inputs and target

Make the PRD, canvas context, target file or page and allowed capability boundary explicit, so success on the wrong target cannot pass as delivery.

02

Execute and retain native output

Generate design through the normal agent and tool path while preserving native nodes, structural snapshots and readable results rather than only a rendered image.

03

Accept with combined evidence

Use target identity, editable structure, screenshots, text and geometry to separately assess delivery validity, evidence completeness and requirement fidelity.

04

Diagnose and regress

Freeze inputs, artifacts, checks and issues so later fixes can be verified under comparable conditions instead of overwriting failures with a best run.

Architecture

Inside the workflow

From requirements to repeatable delivery

Keep inputs, artifacts, evidence and verdicts separate so repairs can be verified against the same task.

Same input · verify again01Freeze inputsPRD / CONTEXT02ExecuteAGENT / TOOLS03Collect evidenceSTRUCTURE / RENDER04ReviewVALIDITY / QUALITY05Diagnose & repairFINDINGS / FIX06RegressionSAME CASE / RECHECK
01 / 06

Freeze inputs

Record requirements, target identity and tool boundaries under one acceptance contract.

Input
Requirements and canvas context
Output
Reproducible task inputs

Public methodology illustration; missing evidence remains unknown.

Technical details

Tri-layer Evaluation & Acceptance Matrix

Delivery Validity

Baseline Feasibility

Verifying whether the artifact actually exists and is usable. Never treat a tool HTTP 200 as a delivery verdict.

  • Target file & page identity check
  • Non-empty nodes & valid geometry
  • Native editable text & layer hierarchy
Evidence Completeness

Auditable Trail

Every verdict must carry reproducible evidence. If evidence is lacking, uncertainty remains explicitly 'Unknown'.

  • Full-canvas visual screenshots
  • JSON structural snapshots
  • Session logs, configs and execution traces
Requirement Fidelity

Specification Adherence

Reviewing artifacts against PRD constraints and visual requirements, then passing actionable findings to diagnosis and repair.

  • Point-by-point PRD check
  • Visual readability & hierarchy
  • Product verdicts independent of diagnosis

Nonempty is not complete

Acceptance begins with target identity and native structure: the correct file or page, visible nonempty content, decodable screenshots, and editable text or design nodes. Only after this evidence exists can requirement and visual judgment begin.

Keep uncertainty explicit

Structural facts, visual readability and business semantics need different evidence. When evidence is missing, the result is unknown, not a guess based on a screenshot or a single success signal. This boundary makes evaluation more trustworthy and diagnosis actionable.

Preserve evidence for iteration

Each run retains its inputs, target identity, native structure, screenshots, sessions and version information. That separates generation failure, delivery failure, incomplete collection and unmet requirements, while enabling comparable regression checks after a fix.

Use findings to drive repair

Each finding retains a location, relevant constraints and direct evidence. Execution records help diagnose the generation process while the original product verdict remains intact. After repair, the same task is reviewed again for the original issue and affected content, producing new regression evidence.

Delivery validity, evidence coverage and product quality are reported separately, with observed issues retained as actionable repair items.

Why it matters

This method connects verifiable targets, native artifacts and regression evidence to determine whether a design was delivered.

Lessons learned

Record execution, evidence, requirements and diagnosis separately; leave unsupported conclusions unknown.