A physician-built lab for evaluating clinical AI across three eval frameworks: single-turn decision support, modular pediatric triage, and multi-turn conversational health agents. Hand-curated traps, blinded results, real failure modes. Built and run from a working pediatric practice.
Clinical AI doesn't have one shape. Each of these workstreams is a separate harness with its own rubric stack, scored independently, because a system that wins on one shape can lose on another. The lab below runs all three.
Accuracy, completeness, and specificity scored on production-relevant primary-care queries. Dual rubric: full CDS and content-only, so packaging can't carry the score.
Two separate suites — known clinical traps (contraindicated drugs, dangerous doses, false premises) and fabricated entities (invented trial names, fake guidelines). Pass rates, not just averages.
Thirty queries whose correct answer changed in the last 24 months. Tests parametric staleness in models without retrieval and stale retrieval in models with it.
Sanitized for public release — exact prompt text stays in the harness so vendors can't train against it. What follows is the shape of each probe and the failure mode it's designed to surface.
Nine systems, scored on the Golden suite with the production rubric. Vendor identities are withheld; categories are accurate. Detailed vendor-attributed reports are shared privately with each team.
| System | Accuracy | Compl. | Specificity | Citations | Latency | $ / query |
|---|
Per-suite breakdown
The framework
Four modules, scored by four specialized judges
What the modules surfaced
Methodology
Forty personas, twenty of them adversarial
What multi-turn evaluation showed
Findings are conclusions about specific systems. Contributions are reusable methods. These are the patterns from this work I think generalize past the specific vendors and models on the leaderboard.