Clinical Decision Support · Independent Evaluation

I measure whether clinical AI is actually safe to put in front of patients and clinicians.

A physician-built lab for evaluating clinical AI across three eval frameworks: single-turn decision support, modular pediatric triage, and multi-turn conversational health agents. Hand-curated traps, blinded results, real failure modes. Built and run from a working pediatric practice.

01

Three eval frameworks — different shapes, different rubrics

Clinical AI doesn't have one shape. Each of these workstreams is a separate harness with its own rubric stack, scored independently, because a system that wins on one shape can lose on another. The lab below runs all three.

01 · 1

What I evaluate at the CDS layer

Clinical content quality

Accuracy, completeness, and specificity scored on production-relevant primary-care queries. Dual rubric: full CDS and content-only, so packaging can't carry the score.

Hallucination safety

Two separate suites — known clinical traps (contraindicated drugs, dangerous doses, false premises) and fabricated entities (invented trial names, fake guidelines). Pass rates, not just averages.

Freshness

Thirty queries whose correct answer changed in the last 24 months. Tests parametric staleness in models without retrieval and stale retrieval in models with it.

What this site covers

01 · 2

The benchmark stack — seven hand-curated suites

01 · 3

Example queries

Sanitized for public release — exact prompt text stays in the harness so vendors can't train against it. What follows is the shape of each probe and the failure mode it's designed to surface.

01 · 4

Blinded leaderboard

Nine systems, scored on the Golden suite with the production rubric. Vendor identities are withheld; categories are accurate. Detailed vendor-attributed reports are shared privately with each team.

System Accuracy Compl. Specificity Citations Latency $ / query
All values 0–100 unless marked. Sorted by content-dimension total.

Per-suite breakdown

How each system holds up across the rest of the stack

01 · 5

How the CDS harness evolved — four months of iteration

01 · 6

Where the eval budget actually goes

01 · 7

What the CDS evaluations showed

02

Pediatric triage assistant — modular by capability

The framework

Four modules, scored by four specialized judges

What the modules surfaced

03

Conversational health agent — scored across twenty-five turns

Methodology

Forty personas, twenty of them adversarial

What multi-turn evaluation showed

04

Methodology contributions — patterns worth taking elsewhere

Findings are conclusions about specific systems. Contributions are reusable methods. These are the patterns from this work I think generalize past the specific vendors and models on the leaderboard.

About the operator