This is a technology demonstration, not a medical device
Synthetic data only. No real patients, no real prescriptions, and nothing here is clinical advice. The catalogues are public reference data and the gate is a software invariant, not a regulatory approval.
This project quotes ISMP, the FDA, the Joint Commission and 21 CFR. It is not affiliated with, endorsed by, or reviewed by any of them. Those citations establish that read-back is an existing requirement; they establish nothing about this software.
Do not enter real patient data into this application.
Every figure on this page carries the command that produced it and the size of the set it came from. A number without a method is not published here, so a figure reads as a dash rather than being filled with an estimate.
Why a figure can read as a dash
A dash has two causes and the command column says which: a named target means the run costs credit and has not been spent, and no command yet means nothing computes the figure.
The pair rule stops every seeded mishearing, and puts its longer question to 21 of 59 correct names
The catch and its cost at equal weight, and each mechanism with its own number. Showing only one of the two would make the metric one-sided.
What the pair rule catches
Written without the pair ruleseeded pair mishearings, each answered by a reflex yes
20/20
Written with the pair rule
0/20
Both arms read every drug name back and differ by the pair rule alone; how often a real caller answers a plain read-back by reflex is not measured. npx tsx scripts/measure/ab-gate.tsn = 20, text candidates
What caught the real recognizer errors: the catalogue check
Recorded recognizer errors the catalogue check refused
21 of 21
Each heard name matches no prescription product. None was heard as a published partner, so the pair rule caught none of them, and its catch above rests on the seeded pairs.
Of the same errors, at or above the drug-name threshold
4 of 21
A threshold alone would have written every one of them.
make coverage-matrixn = 21, synthesised speech through the live recognizer
What the pair rule costs on correct names
Correct drug names put to the pair rule's longer, contrastive question
21 of 5935.6% [24.6%, 48.3%]
The shipped gate as a whole, not the pair rule alone, asks about every correct drug name: 59 of 59.
25plain read-back
13threshold re-ask below 0.95
21pair rule: contrastive question
With the pair rule switched off, the threshold would take 16 of the 59. make coverage-matrixn = 59, recorded confidences from synthesised speech through the live recognizer, scored against the full 2023 ISMP list
How the cost is counted
Every one of the 59 drug names was heard correctly, and each is read back once by policy; each mechanism pays for its own share, and the pair rule’s share is the names on the ISMP list.
A contrastive question runs about 11.8 s (28.7 words) against about 2.5 s (6.0 words) for a plain read-back: 9.4 s more per name asked that way. The seconds use 2.43 words per second, the desktop synthesiser's rate over eval/control; the agent's own voice has not been timed, and each spelled letter counts as a word, so the contrastive figure overstates the letters. npx tsx scripts/measure/coverage-matrix.tsn = 59 plain, 21 contrastive, synthesised speech
Published totals match raw runs, and a test fails if they disagree: tests/features/metrics/raw-run-agreement.test.ts
What the shipped gate stops, and what it asks
Standing read-back, threshold and contrastive pair rule, as the server publishes them. Brackets are 95% Wilson intervals.
One row per figure. The input column says what went in: a live socket, synthesised speech (TTS) sent through the live recognizer, text with assigned values, or a synthesised fixture. A dash means not measured, never zero.
Figure
Value
Input
Command
n
Recognizer errors stopped before any plain read-back, shipped policyRecorded confidences from synthesised speech through the live recognizer, replayed through the shipped gate: standing read-back, threshold and pair rule together.
100.0% [84.5%, 100.0%]
TTS
npx tsx scripts/measure/coverage-matrix.ts
21the control and native 16 kHz runs
Correct drug names the shipped gate asks about, the standing read-back included
59/59
TTS
npx tsx scripts/measure/coverage-matrix.ts
59the control and native 16 kHz runs
Correct drug names given a contrastive question under the full ISMP list
35.6% [24.6%, 48.3%]
TTS
npx tsx scripts/measure/coverage-matrix.ts
59the control and native 16 kHz runs
Pair mishearings a reflex yes would write, shipped policyBoth arms read every drug name back and differ by the pair rule alone; a caller's reflex yes is assumed, because how often a real caller gives one is not measured.
0/20
text
npx tsx scripts/measure/ab-gate.ts
20seeded text candidates misheard inside a curated pair
Pair mishearings a reflex yes would write, same policy without the pair rule
20/20
text
npx tsx scripts/measure/ab-gate.ts
20seeded text candidates misheard inside a curated pair
Distinct pairs in the full 2023 ISMP list, the product ruleCounted from the full 2023 ISMP list, which the product applies, and the built catalogue. The 20 curated pairs are the evaluation core, not the rule.
514
text
npx tsx scripts/measure/ismp-coverage.ts
514the parsed list and the built catalogue
Catalogue drugs carrying a name on the ISMP list
502 of 3730 (13.5%)
text
npx tsx scripts/measure/ismp-coverage.ts
3730the parsed list and the built catalogue
The held-out set: sealed, opened once, and one hypothesis failed
Thresholds are chosen defaults, not tuned on any set: the drug-name threshold is 0.95 and the strength threshold 0.92, reasoned from the cost of an error before any audio existed. The held-out set was drawn and sealed with its hypothesis written down first, then opened once.
The pre-registered hypothesis did not replicate. It predicted that rarer names fail more, with intervals clear of each other; the three strata overlap, and the result is published as a negative one rather than dropped.
What replicated is the overall error rate: 26.7% on the held-out set against 27.5% on the control corpus.
Rare stratum30.0% [14.5%, 51.9%]
Middle stratum30.0% [14.5%, 51.9%]
Common stratum20.0% [8.1%, 41.6%]
npx tsx scripts/measure/analyse-rarity.ts --set eval/heldout --strata 3n = 20 each, three rarity strata of eval/heldout
The held-out run, as the report publishes it, each row with the command that reprints it from the recorded file.
Figure
Value
Input
Command
n
Entity error rate on the held-out setSixty names never measured before, drawn and sealed before the run, then opened once. It sits beside 27.5% on the control corpus, so the recognizer's difficulty with rare names generalises to names it had never been measured on.
26.7% [17.1%, 39.0%]
TTS
npx tsx scripts/eer/report.ts eval/heldout
60eval/heldout, recorded by make eval-heldout
Rare stratum: at most one catalogue combinationThe pre-registered hypothesis predicted this stratum would fail most, with an interval clear of the common one.
Common stratum: five or more combinationsThe interval overlaps the rare one heavily, so the effect is not demonstrated and the hypothesis is published as not supported.
Of the 16 held-out errors, 4 sat at or above the 0.95 threshold. Two are typos our own sampler drew from the FDA file (llevofloxacin and epineprine), so the recognizer was scored wrong for hearing the real word; the other two, oteseconazole and chlorthalidone, are genuine recognizer errors that a threshold alone would have passed. npx tsx scripts/eer/report.ts eval/heldout prints all 16. Without the two typo items the rate is 24.1% [15.0%, 36.5%], n = 58; both figures are published, because choosing the flattering one after seeing them is what the seal exists to prevent.
The gate’s own catch and false-ask rates on the held-out set are not scored yet; they read as dashes under Not measured yet.
Every other figure, and what the gate would cost an order
The full benchmark and the cost readings each have their own page.