Readback
Source code on GitHub (opens in a new tab)

Measurements

Every figure on this page carries the command that produced it and the size of the set it came from. A number without a method is not published here, so a figure reads as a dash rather than being filled with an estimate.

Why a figure can read as a dash

A dash has two causes and the command column says which: a named target means the run costs credit and has not been spent, and no command yet means nothing computes the figure.

The pair rule stops every seeded mishearing, and puts its longer question to 21 of 59 correct names

The catch and its cost at equal weight, and each mechanism with its own number. Showing only one of the two would make the metric one-sided.

What the pair rule catches

Written without the pair ruleseeded pair mishearings, each answered by a reflex yes
20/20
Written with the pair rule
0/20

Both arms read every drug name back and differ by the pair rule alone; how often a real caller answers a plain read-back by reflex is not measured. npx tsx scripts/measure/ab-gate.ts n = 20, text candidates

What caught the real recognizer errors: the catalogue check

Recorded recognizer errors the catalogue check refused
21 of 21
Each heard name matches no prescription product. None was heard as a published partner, so the pair rule caught none of them, and its catch above rests on the seeded pairs.
Of the same errors, at or above the drug-name threshold
4 of 21
A threshold alone would have written every one of them.
make coverage-matrix n = 21, synthesised speech through the live recognizer

What the pair rule costs on correct names

Correct drug names put to the pair rule's longer, contrastive question
21 of 5935.6% [24.6%, 48.3%]

The shipped gate as a whole, not the pair rule alone, asks about every correct drug name: 59 of 59.

  • 25plain read-back
  • 13threshold re-ask below 0.95
  • 21pair rule: contrastive question

With the pair rule switched off, the threshold would take 16 of the 59. make coverage-matrix n = 59, recorded confidences from synthesised speech through the live recognizer, scored against the full 2023 ISMP list

How the cost is counted

Every one of the 59 drug names was heard correctly, and each is read back once by policy; each mechanism pays for its own share, and the pair rule’s share is the names on the ISMP list.

A contrastive question runs about 11.8 s (28.7 words) against about 2.5 s (6.0 words) for a plain read-back: 9.4 s more per name asked that way. The seconds use 2.43 words per second, the desktop synthesiser's rate over eval/control; the agent's own voice has not been timed, and each spelled letter counts as a word, so the contrastive figure overstates the letters. npx tsx scripts/measure/coverage-matrix.ts n = 59 plain, 21 contrastive, synthesised speech

Published totals match raw runs, and a test fails if they disagree: tests/features/metrics/raw-run-agreement.test.ts

What the shipped gate stops, and what it asks

Standing read-back, threshold and contrastive pair rule, as the server publishes them. Brackets are 95% Wilson intervals.

One row per figure. The input column says what went in: a live socket, synthesised speech (TTS) sent through the live recognizer, text with assigned values, or a synthesised fixture. A dash means not measured, never zero.
FigureValueInputCommandn
Recognizer errors stopped before any plain read-back, shipped policyRecorded confidences from synthesised speech through the live recognizer, replayed through the shipped gate: standing read-back, threshold and pair rule together.100.0% [84.5%, 100.0%]TTSnpx tsx scripts/measure/coverage-matrix.ts21the control and native 16 kHz runs
Correct drug names the shipped gate asks about, the standing read-back included59/59TTSnpx tsx scripts/measure/coverage-matrix.ts59the control and native 16 kHz runs
Correct drug names given a contrastive question under the full ISMP list35.6% [24.6%, 48.3%]TTSnpx tsx scripts/measure/coverage-matrix.ts59the control and native 16 kHz runs
Pair mishearings a reflex yes would write, shipped policyBoth arms read every drug name back and differ by the pair rule alone; a caller's reflex yes is assumed, because how often a real caller gives one is not measured.0/20textnpx tsx scripts/measure/ab-gate.ts20seeded text candidates misheard inside a curated pair
Pair mishearings a reflex yes would write, same policy without the pair rule20/20textnpx tsx scripts/measure/ab-gate.ts20seeded text candidates misheard inside a curated pair
Distinct pairs in the full 2023 ISMP list, the product ruleCounted from the full 2023 ISMP list, which the product applies, and the built catalogue. The 20 curated pairs are the evaluation core, not the rule.514textnpx tsx scripts/measure/ismp-coverage.ts514the parsed list and the built catalogue
Catalogue drugs carrying a name on the ISMP list502 of 3730 (13.5%)textnpx tsx scripts/measure/ismp-coverage.ts3730the parsed list and the built catalogue

The held-out set: sealed, opened once, and one hypothesis failed

Thresholds are chosen defaults, not tuned on any set: the drug-name threshold is 0.95 and the strength threshold 0.92, reasoned from the cost of an error before any audio existed. The held-out set was drawn and sealed with its hypothesis written down first, then opened once.

The pre-registered hypothesis did not replicate. It predicted that rarer names fail more, with intervals clear of each other; the three strata overlap, and the result is published as a negative one rather than dropped.

What replicated is the overall error rate: 26.7% on the held-out set against 27.5% on the control corpus.

  • Rare stratum30.0% [14.5%, 51.9%]
  • Middle stratum30.0% [14.5%, 51.9%]
  • Common stratum20.0% [8.1%, 41.6%]

npx tsx scripts/measure/analyse-rarity.ts --set eval/heldout --strata 3 n = 20 each, three rarity strata of eval/heldout

The held-out run, as the report publishes it, each row with the command that reprints it from the recorded file.
FigureValueInputCommandn
Entity error rate on the held-out setSixty names never measured before, drawn and sealed before the run, then opened once. It sits beside 27.5% on the control corpus, so the recognizer's difficulty with rare names generalises to names it had never been measured on.26.7% [17.1%, 39.0%]TTSnpx tsx scripts/eer/report.ts eval/heldout60eval/heldout, recorded by make eval-heldout
Rare stratum: at most one catalogue combinationThe pre-registered hypothesis predicted this stratum would fail most, with an interval clear of the common one.30.0% [14.5%, 51.9%]TTSnpx tsx scripts/measure/analyse-rarity.ts --set eval/heldout --strata 320eval/heldout, recorded by make eval-heldout
Middle stratum: two to four combinationsIdentical to the rare stratum, where the hypothesis predicted a decline.30.0% [14.5%, 51.9%]TTSnpx tsx scripts/measure/analyse-rarity.ts --set eval/heldout --strata 320eval/heldout, recorded by make eval-heldout
Common stratum: five or more combinationsThe interval overlaps the rare one heavily, so the effect is not demonstrated and the hypothesis is published as not supported.20.0% [8.1%, 41.6%]TTSnpx tsx scripts/measure/analyse-rarity.ts --set eval/heldout --strata 320eval/heldout, recorded by make eval-heldout
The 4 errors at or above the threshold

Of the 16 held-out errors, 4 sat at or above the 0.95 threshold. Two are typos our own sampler drew from the FDA file (llevofloxacin and epineprine), so the recognizer was scored wrong for hearing the real word; the other two, oteseconazole and chlorthalidone, are genuine recognizer errors that a threshold alone would have passed. npx tsx scripts/eer/report.ts eval/heldout prints all 16. Without the two typo items the rate is 24.1% [15.0%, 36.5%], n = 58; both figures are published, because choosing the flattering one after seeing them is what the seal exists to prevent.

The gate’s own catch and false-ask rates on the held-out set are not scored yet; they read as dashes under Not measured yet.

Every other figure, and what the gate would cost an order

The full benchmark and the cost readings each have their own page.