95%+
Core Acceptance Coverage
20-30m → 3-5m979
Atomic Business Criteria
84 tasks / 160 pages95%+
Complex Edit Success Rate
-60% tool calls95%+
MCP Generation E2E Rate
-50% concurrency errorsBackground
A design agent does not produce a single text answer. It must turn a PRD and canvas context into design content that another person can continue to edit. A successful tool call, a screenshot or a nonempty node is not proof of completed delivery.
Real delivery answers three different questions: is the target correct, is the artifact usable with complete evidence, and were requirements satisfied? Collapsing them into one success state hides failures and provides no useful iteration path.
Solution
Freeze inputs and target
Make the PRD, canvas context, target file or page and allowed capability boundary explicit, so success on the wrong target cannot pass as delivery.
Execute and retain native output
Generate design through the normal agent and tool path while preserving native nodes, structural snapshots and readable results rather than only a rendered image.
Accept with combined evidence
Use target identity, editable structure, screenshots, text and geometry to separately assess delivery validity, evidence completeness and requirement fidelity.
Diagnose and regress
Freeze inputs, artifacts, checks and issues so later fixes can be verified under comparable conditions instead of overwriting failures with a best run.
Architecture
From requirements to repeatable delivery
Keep inputs, artifacts, evidence and verdicts separate so repairs can be verified against the same task.
Freeze inputs
Record requirements, target identity and tool boundaries under one acceptance contract.
- Input
- Requirements and canvas context
- Output
- Reproducible task inputs
Public methodology illustration; missing evidence remains unknown.
Technical details
Tri-layer Evaluation & Acceptance Matrix
Delivery Validity
Baseline Feasibility
Verifying whether the artifact actually exists and is usable. Never treat a tool HTTP 200 as a delivery verdict.
- Target file & page identity check
- Non-empty nodes & valid geometry
- Native editable text & layer hierarchy
Evidence Completeness
Auditable Trail
Every verdict must carry reproducible evidence. If evidence is lacking, uncertainty remains explicitly 'Unknown'.
- Full-canvas visual screenshots
- JSON structural snapshots
- Session logs, configs and execution traces
Requirement Fidelity
Specification Adherence
Reviewing artifacts against PRD constraints and visual requirements, then passing actionable findings to diagnosis and repair.
- Point-by-point PRD check
- Visual readability & hierarchy
- Product verdicts independent of diagnosis
Nonempty is not complete
Acceptance begins with target identity and native structure: the correct file or page, visible nonempty content, decodable screenshots, and editable text or design nodes. Only after this evidence exists can requirement and visual judgment begin.
Keep uncertainty explicit
Structural facts, visual readability and business semantics need different evidence. When evidence is missing, the result is unknown, not a guess based on a screenshot or a single success signal. This boundary makes evaluation more trustworthy and diagnosis actionable.
Preserve evidence for iteration
Each run retains its inputs, target identity, native structure, screenshots, sessions and version information. That separates generation failure, delivery failure, incomplete collection and unmet requirements, while enabling comparable regression checks after a fix.
Use findings to drive repair
Each finding retains a location, relevant constraints and direct evidence. Execution records help diagnose the generation process while the original product verdict remains intact. After repair, the same task is reviewed again for the original issue and affected content, producing new regression evidence.
Delivery validity, evidence coverage and product quality are reported separately, with observed issues retained as actionable repair items.
Why it matters
This method connects verifiable targets, native artifacts and regression evidence to determine whether a design was delivered.
Lessons learned
Record execution, evidence, requirements and diagnosis separately; leave unsupported conclusions unknown.