Readback
Source code on GitHub (opens in a new tab)

Benchmark

One row per figure. The input column says what went in: a live socket, synthesised speech (TTS) sent through the live recognizer, text with assigned values, or a synthesised fixture. A dash means not measured, never zero.

What the recognizer got wrong, and how confident it was

Recorded runs of synthesised speech through the live recognizer, scored for the drug name and replayed through the real validator and gate. Brackets are 95% Wilson intervals.

Measured figures, each with its command and set size.
FigureValueInputCommandn
Entity error rate, development setHow often the recognizer returned a different drug name than the one spoken. 40 drug names spoken by two synthetic voices; the development set, on which no threshold was tuned.0.0% [0.0, 8.8]TTSmake eval40utterances in eval/dev
Entity error rate, control set, resampled40 rarer names chosen before measuring, captured at 48 kHz and resampled to 16 kHz.27.5% [16.1, 42.8]TTSmake eval-control40utterances in eval/control
Entity error rate, control set, native 16 kHzThe same 40 names synthesised at 16 kHz, with no resampling in the path.25.0% [14.2, 40.2]TTSmake eval-native1640utterances in eval/native16
Errors the recognizer reported at or above the thresholdRecognizer errors whose own reported confidence sat at or above the drug-name threshold. A confidence check alone would have written every one of these into an order, which is the reason this product does not rely on one.4 of 21TTSmake coverage-matrix21recorded errors across the control and native 16 kHz runs
Correct values the threshold would re-askValues the recognizer got right but reported below the threshold, counted with the pair rule switched off. The shipped policy reads every drug name back once regardless, so the threshold changes which question is asked, not whether one is; the catalogue check adds none of its own.16 of 59TTSmake coverage-matrix59correct values across the control and native 16 kHz runs
Errors the catalogue refusesRecorded errors that no prescription product in the built catalogue matches, so the gate refuses them regardless of confidence. Measured by running the real validator and the real gate over each recorded utterance.21 of 21TTSmake coverage-matrix21recorded errors, assigned in the gate's own branch order

Not measured yet, and what would measure each one

Each of these reads as a dash. A named target means the run costs credit and has not been spent; no command yet means nothing computes the figure.

Figures not measured yet, with what would produce each one.
FigureValueInputCommandn
LASA catch rateOf the held-out items where the recognizer returned the wrong member of a published pair, the share the gate stopped before the value entered the order.—TTSno command yet—the held-out set; no script scores the gate on it yet
False-ask rateHow often the gate asked again when the value was already correct. This is the cost side of the idea and the only honest answer to whether the agent re-asks constantly.—TTSno command yet—the held-out set; no script scores the gate on it yet
Accepted-wrong countValues the gate let through that were in fact wrong. A single one is a safety failure and raises the threshold for that field.—TTSno command yet—the held-out set; no script scores the gate on it yet
Read-back match rateShare of requested read-backs the caller confirmed on the first attempt.—live socketno command yet—live sessions with a caller; none has been recorded
Escalation rateShare of sessions where a critical field exhausted its attempts, so the order was refused and marked as needing a pharmacist.—live socketno command yet—live sessions with a caller; none has been recorded
Words re-said per orderHow many words the caller had to say a second time, per finished order, because the gate asked again. The cost of the idea in the caller's own breath.—live socketno command yet—finished live orders; none has been recorded
Word to gate decisionMilliseconds from the last word of the source span to the gate decision being rendered. This is a browser measurement: word-level timings never reach our server, so it is labelled as client-side and not as a server figure.—live socketmake measure—each live run, spaced about 24 seconds apart
Time to first audioServer-reported milliseconds before the agent's first audio frame, read from GET /v1/sessions/{id}. This one is honest turn-level data from AssemblyAI rather than a browser stopwatch.—live socketmake measure—each live run, spaced about 24 seconds apart

Checksums, calibration and rarity

Published totals match raw runs: tests/stats/benchmark-agreement.test.ts fails when a figure here is absent from eval/REPORT.md or its command is not a step that make honest re-runs.

Checksum audits, calibration, rarity and human-voice rows the server publishes beside the shipped policy.
FigureValueInputCommandn
Entity error rate of the recognizer, 95% WilsonThe control-set run from the table above, as the report prints it: the same measurement, not a second one.27.5%TTSnpx tsx scripts/eer/report.ts eval/control40
NPI single-digit substitutions caught by the checksumEvery single-digit substitution and every adjacent transposition of 200 valid identifiers of each kind, enumerated rather than sampled. The two checksums are not equally strong, and the rows say by how much.100.0%textnpx tsx scripts/measure/audit-checksums.ts 20018000
NPI adjacent transpositions caught by the checksum97.9%textnpx tsx scripts/measure/audit-checksums.ts 200840
DEA single-digit substitutions caught by the checksum95.2%textnpx tsx scripts/measure/audit-checksums.ts 20012600
Observed accuracy at reported confidence 0.95 to 0.99Of the recorded utterances whose reported confidence fell in this band, the share whose drug name was right.89.7% [76.4%, 95.9%]TTSnpx tsx scripts/measure/analyse-calibration.ts39
Error rate on drugs with at most one catalogue combinationThe control corpus split by catalogue combinations. The pre-registered held-out run did not replicate this gap; the Measurements page publishes that negative result.43.8%TTSnpx tsx scripts/measure/analyse-rarity.ts16
Entity error rate on human voices—live socketmake eval-live—
False confirmations on non-commands, human voices—live socketmake eval-live—
Recognizer drift toward a hinted LASA name (keyterms ablation)—live socketnpx tsx scripts/measure/keyterms-ablation.ts—
End of speech to gate decision, browser—live socketnot measured—