All projects
AI / Evaluation

Sapien · Proof of Quality

Evaluation layer for AI training data and agent output

roleSenior Product Manager
timelineOct 2025 – Present
focusProduct Strategy, Design Partnerships, AI Evaluation
stackGo, React 19, LLM Judges, Human Expert Panels, Base

About

Proof of Quality is Sapien's evaluation layer for AI work: a customer defines quality as a rubric, a panel of LLM judges and human experts scores the work against it, and the output is an auditable quality report with 2 scores on a 0 to 100 scale that the customer's own clients can trust. The commercial idea is simple: pay for verified quality, not review hours.

I own strategy, roadmap, and success metrics, and I ship inside the product itself. That means writing the launch specs, running the pilots that test the thesis with real enterprise partners, redesigning the surfaces raters use every day, and building the demo tooling that takes a cold prospect to a live evaluation project in 1 session.

Process

1

Product definition

DRI on 2 of the P1 scoping documents for V1: customer onboarding (introduced the demo project as the onboarding mechanism) and rater onboarding and credentialing. Ran a critical review of the V1 spec that resolved 5 launch-blocking defects before engineering kickoff, including a quality metric defined 2 incompatible ways and AI review panels copying each other's votes. Wrote the product vocabulary and rubric guardrails engineering builds against.

2

Design-partner trials

Ran Sapien's most active design partnership end to end: a trial that re-evaluated a security firm's AI Solidity auditor against 5 senior human auditors. The panel flagged 21 of 40 findings as false positives in 77 minutes, with an auditable verdict on 98%. The Phase 2 partnership proposal that followed was approved.

3

Shipping in-product

Redesigned the rater results experience around a verdict-first layout (~450 lines of new React components with Storybook and Vitest coverage, shipped through review across 2 pull requests). Built an automation skill that turns a company name into a fully scored demo project: researched dataset, authored rubric, LLM judge panel, packaged quality report. A 28-item demo scores in minutes.

4

Forward-deployed GTM

Ran 120+ customer discovery and go-to-market conversations, producing 3 pilot proposals for security firms including CertiK and Sherlock and a converted design partner whose CEO chose Sapien over extending his existing tooling after a live demo. Worked 9 side events at Consensus Miami 2026: 23 contacts, 12 qualified leads, 8 rated high-interest, and the field report that reshaped our messaging.

Outcome

2 completed evaluation trials on real production AI output, an approved Phase 2 enterprise partnership, a pilot pipeline across the Web3 security vertical, and pay-for-the-proof pricing ($1,000 per project, per-event fees, $500 per report) pressure-tested in customer conversations. The trial data confirmed the core thesis: consensus sharpens as the review panel grows.