Agent Eval & Data Flywheel
Led 84-task / 160-page / 979-atomic-criteria real-world benchmark; combined native Figma parsing, screenshots, and Rule/LLM-as-a-Judge with automated trace root-cause attribution, achieving 95%+ core acceptance coverage.
Main Site Technology (Operations Tech) · Agent Developer · 2026.08–Present
Leading evaluation benchmarks, Figma JSX declarative structured editing, and MCP integration for M0 Design (PRD → Figma → Code) Agent: connecting generation, evaluation, and production invocation into a closed engineering loop.
Led 84-task / 160-page / 979-atomic-criteria real-world benchmark; combined native Figma parsing, screenshots, and Rule/LLM-as-a-Judge with automated trace root-cause attribution, achieving 95%+ core acceptance coverage.
Serialized Figma SceneGraph into JSX structures for model read/write with safe Reconcile diff writebacks; 95%+ complex design edit success rate on 143+ test cases, reducing fine-grained tool calls by 60%+.
Packaged generation capabilities into MCP for internal agents (Agent → MCP → Figma Runtime / Host); resolved write timeouts and runtime crashes, raising complex generation E2E success rate to 95%+ and halving concurrency errors.
An output needs a clear target identity, editable structure and constraints; a successful tool call is not a delivery verdict.
Figma AST snapshots, visual screenshots, and semantic requirements each have distinct evidence sources. When evidence is insufficient, uncertainty stays explicit.
Build a continuous loop: Online Cases → Eval → Root-cause Attribution → Fixes Merged → Regression → Dataset Re-injection.
Read the engineering methodology on verifiable agent delivery