Benchmark
One row per figure. The input column says what went in: a live socket, synthesised speech (TTS) sent through the live recognizer, text with assigned values, or a synthesised fixture. A dash means not measured, never zero.
What the recognizer got wrong, and how confident it was
Recorded runs of synthesised speech through the live recognizer, scored for the drug name and replayed through the real validator and gate. Brackets are 95% Wilson intervals.
| Figure | Value | Input | Command | n |
|---|---|---|---|---|
| Entity error rate, development setHow often the recognizer returned a different drug name than the one spoken. 40 drug names spoken by two synthetic voices; the development set, on which no threshold was tuned. | 0.0% [0.0, 8.8] | TTS | make eval | 40utterances in eval/dev |
| Entity error rate, control set, resampled40 rarer names chosen before measuring, captured at 48 kHz and resampled to 16 kHz. | 27.5% [16.1, 42.8] | TTS | make eval-control | 40utterances in eval/control |
| Entity error rate, control set, native 16 kHzThe same 40 names synthesised at 16 kHz, with no resampling in the path. | 25.0% [14.2, 40.2] | TTS | make eval-native16 | 40utterances in eval/native16 |
| Errors the recognizer reported at or above the thresholdRecognizer errors whose own reported confidence sat at or above the drug-name threshold. A confidence check alone would have written every one of these into an order, which is the reason this product does not rely on one. | 4 of 21 | TTS | make coverage-matrix | 21recorded errors across the control and native 16 kHz runs |
| Correct values the threshold would re-askValues the recognizer got right but reported below the threshold, counted with the pair rule switched off. The shipped policy reads every drug name back once regardless, so the threshold changes which question is asked, not whether one is; the catalogue check adds none of its own. | 16 of 59 | TTS | make coverage-matrix | 59correct values across the control and native 16 kHz runs |
| Errors the catalogue refusesRecorded errors that no prescription product in the built catalogue matches, so the gate refuses them regardless of confidence. Measured by running the real validator and the real gate over each recorded utterance. | 21 of 21 | TTS | make coverage-matrix | 21recorded errors, assigned in the gate's own branch order |
Not measured yet, and what would measure each one
Each of these reads as a dash. A named target means the run costs credit and has not been spent; no command yet means nothing computes the figure.
| Figure | Value | Input | Command | n |
|---|---|---|---|---|
| LASA catch rateOf the held-out items where the recognizer returned the wrong member of a published pair, the share the gate stopped before the value entered the order. | — | TTS | no command yet | —the held-out set; no script scores the gate on it yet |
| False-ask rateHow often the gate asked again when the value was already correct. This is the cost side of the idea and the only honest answer to whether the agent re-asks constantly. | — | TTS | no command yet | —the held-out set; no script scores the gate on it yet |
| Accepted-wrong countValues the gate let through that were in fact wrong. A single one is a safety failure and raises the threshold for that field. | — | TTS | no command yet | —the held-out set; no script scores the gate on it yet |
| Read-back match rateShare of requested read-backs the caller confirmed on the first attempt. | — | live socket | no command yet | —live sessions with a caller; none has been recorded |
| Escalation rateShare of sessions where a critical field exhausted its attempts, so the order was refused and marked as needing a pharmacist. | — | live socket | no command yet | —live sessions with a caller; none has been recorded |
| Words re-said per orderHow many words the caller had to say a second time, per finished order, because the gate asked again. The cost of the idea in the caller's own breath. | — | live socket | no command yet | —finished live orders; none has been recorded |
| Word to gate decisionMilliseconds from the last word of the source span to the gate decision being rendered. This is a browser measurement: word-level timings never reach our server, so it is labelled as client-side and not as a server figure. | — | live socket | make measure | —each live run, spaced about 24 seconds apart |
| Time to first audioServer-reported milliseconds before the agent's first audio frame, read from GET /v1/sessions/{id}. This one is honest turn-level data from AssemblyAI rather than a browser stopwatch. | — | live socket | make measure | —each live run, spaced about 24 seconds apart |
Checksums, calibration and rarity
Published totals match raw runs: tests/stats/benchmark-agreement.test.ts fails when a figure here is absent from eval/REPORT.md or its command is not a step that make honest re-runs.
| Figure | Value | Input | Command | n |
|---|---|---|---|---|
| Entity error rate of the recognizer, 95% WilsonThe control-set run from the table above, as the report prints it: the same measurement, not a second one. | 27.5% | TTS | npx tsx scripts/eer/report.ts eval/control | 40 |
| NPI single-digit substitutions caught by the checksumEvery single-digit substitution and every adjacent transposition of 200 valid identifiers of each kind, enumerated rather than sampled. The two checksums are not equally strong, and the rows say by how much. | 100.0% | text | npx tsx scripts/measure/audit-checksums.ts 200 | 18000 |
| NPI adjacent transpositions caught by the checksum | 97.9% | text | npx tsx scripts/measure/audit-checksums.ts 200 | 840 |
| DEA single-digit substitutions caught by the checksum | 95.2% | text | npx tsx scripts/measure/audit-checksums.ts 200 | 12600 |
| Observed accuracy at reported confidence 0.95 to 0.99Of the recorded utterances whose reported confidence fell in this band, the share whose drug name was right. | 89.7% [76.4%, 95.9%] | TTS | npx tsx scripts/measure/analyse-calibration.ts | 39 |
| Error rate on drugs with at most one catalogue combinationThe control corpus split by catalogue combinations. The pre-registered held-out run did not replicate this gap; the Measurements page publishes that negative result. | 43.8% | TTS | npx tsx scripts/measure/analyse-rarity.ts | 16 |
| Entity error rate on human voices | — | live socket | make eval-live | — |
| False confirmations on non-commands, human voices | — | live socket | make eval-live | — |
| Recognizer drift toward a hinted LASA name (keyterms ablation) | — | live socket | npx tsx scripts/measure/keyterms-ablation.ts | — |
| End of speech to gate decision, browser | — | live socket | not measured | — |