Limitations
All 14 things this project cannot prove, stated before a reader finds them. Each carries its status: what is measured, what a machine check enforces, what is assumed, and what is false and admitted. The complete account, with the tests that pin each one, is docs/limitations.md in the repository.
The trust boundary sits in the browser
Who can make the gate accept a value it should not, and what we did and did not do about it.
Provenance is computed in the browser
AdmittedFalse against a hostile client, and stated as false
The browser holds the recognizer socket directly, so word timings and per-word certainties never pass through our server; the client posts them to an unauthenticated route. Anyone with DevTools can post arbitrary words, and the gate will accept a value carrying that provenance: it checks that a value traces to words the session reported, not that those words were spoken. The route enforces bounds instead of a secret: a validated session id and caps on words per turn, turns per session and concurrent sessions. Relaying audio through our own host would fix it, at the cost of an always-on process this design deliberately does without.
Two guarantees rest on checks reading the tree correctly
EnforcedEnforced, each exploit kept as a test
The gate-invariant check refuses a ConfirmedValue assertion in either TypeScript syntax outside the gate, any double assertion through unknown, and any type-checker suppression in product code. The secrets check reads the path app/api, not any directory merely named api, and fails when the key is absent from it. Each exploit is kept as a test, and every ratchet has a positive control, because a check that passes when its subject is missing manufactures confidence.
Known security weaknesses left open
AdmittedTen open, each defensible for synthetic data only
Ten known weaknesses of low severity are open.
The ten open weaknesses
- A session id is a bearer capability: its holder can post turns, finalize and mint a reconnect token.
- Looking up an unknown session id costs up to three storage listings, anonymously.
- Finalize trusts an explicit origin label; a missing or unknown one still defaults to live, but the session's holder can set it on purpose.
- A deployment with no tool secret, or one too short, says so in the 401.
- The agent's own hint is quoted back when the gate answers it, as a spoken echo rather than an instruction.
- A store error's text reaches the client; a test pins that no token or URL is included.
- Bodies are parsed before bounds are checked; the platform limits their size.
- The metrics route is uncached, so every view re-reads every stored session.
- Blob objects are public, and the random id in the path is the only secret.
- Six dependency advisories remain: four fixed only by a major upgrade of next or @vercel/blob, not made, and two in Playwright, a test dependency counted through an optional peer of next.
The evidence is synthetic, and one hypothesis failed
Where the published figures stop generalising, and what was deliberately left unmeasured.
A caller who corrects themselves is detected only through six markers
MeasuredMeasured behaviour, pinned by tests; the gap is narrowed, not closed
“Lisinopril, no wait, losartan” refuses lisinopril as a value the caller took back (E_RETRACTED_VALUE, entering the validator branch). The markers are “no wait”, “sorry”, “I mean”, “actually”, “scratch that” and “not X, Y”, with the replacement within four words. A correction without a marker, one spread across two turns, or one using other words such as “rather” or “make that” is not detected; there the read-back is the mitigation.
The rule is the published list; the evaluation is the curated core
Measured
The pair rule applies the full 2023 ISMP List of Confused Drug Names, 514 pairs parsed from the published PDF with the page and row of each kept. Our measured catches and the demo rest on a hand-curated table of 20 pairs, each checked by hand against its row; the rest of the list is applied, not separately evaluated. Matching strips salt forms, because the catalogue stores tramadol hydrochloride where the pair says tramadol. The rule knows only what the list publishes: lisinopril and bisoprolol are on neither tier and get the standing read-back only.
The evaluation corpus is synthesised
MeasuredMeasured against synthetic speech, and labelled everywhere
We found no open English corpus of human speech reading drug names, so every entity error rate is measured against a desktop synthesiser through the live recognizer, and every figure derived from it inherits that boundary. Deliberately not measured:
- Accuracy on human speech: no open corpus exists that we found.
- NPI existence against the live registry: our numbers are synthetic and would only add noise.
- Threshold optima: see the next entry.
- Keyterms that include drug names: it biases the recognizer toward the strings the rules check.
- Market size in money: a figure we cannot source would break the only rule this project is about.
Thresholds are chosen defaults, not measured optima
AssumedAssumption, stated at the head of the report
Each threshold was reasoned from the cost of an error in its field before any audio existed and was not tuned on any set, the held-out set included. Whether a slightly different threshold would do better is not measured.
Our own pre-registered hypothesis failed
AdmittedNot supported, published anyway
We predicted that rarer drug names would be measurably harder to recognise, and wrote the rule down before the audio existed. On the sealed held-out set the rare and middle strata are identical and every interval overlaps. The negative result stays published; what replicated is the overall error rate the product rests on.
Word-to-gate latency is a browser measurement
DisclosedLabelled as client-side wherever it appears
The session endpoint returns time to first audio and tool timings, but not word-level timings, so the time from a word to the gate's decision can only be measured in the browser. It is not measured yet.
Running it is not clinical use
What the sockets tell us, where the audio goes, and what a caller must never rely on this for.
Close codes are observations, not specification
DisclosedObserved; the vendor documents none
AssemblyAI documents no WebSocket close codes at all. We measured 1000 and 1008 ourselves, and 1006 once; 1008 is what the rate limiter actually sends, and we have never observed the 3009 its condition is supposed to produce. 3006 comes from another team's measurement, and 3007, 3008 and 3009 from vendor prose.
Voice audio leaves this application for a third-party vendor
DisclosedDisclosed, not independently verified
Both sockets stream the caller's voice to AssemblyAI: the recognizer to its global endpoint and the voice agent to its US region, pinned because stored agents live in one region and the global host sent a US server and a European browser to different stores. No data-residency choice was made or evaluated; the EU endpoints are not used. What the vendor retains, and for how long, is governed by its own terms, which we have not summarised because a summary we had not verified would be an unsourced claim. No consent screen names the vendor before a session starts. Every demonstration and evaluation run in this project used synthetic speech.
In a real emergency, do not use this application
DisclosedDisclaimer, with an action
If something is wrong right now, such as an allergic reaction, a medication error already taken, or any symptom that feels like an emergency, call 911, or 988 for a mental health crisis, immediately. This application does not call emergency services, does not triage symptoms, and routes nothing to a human faster than a phone would.
This is not a medical device
DisclosedStated on every page
A technology demonstration on synthetic data: no real patients, no real prescriptions, no clinical use and no claim of regulatory approval or review. The project quotes ISMP, the FDA, the Joint Commission and 21 CFR and is not affiliated with, endorsed by or reviewed by any of them. Do not enter real patient data.
The regulatory citations carry no penalty figure
DisclosedCited by clause number; no sanction figure sourced
Read-back is cited to ICAO Annex 11, the Joint Commission's verbal-order goal and 21 CFR 1306.12(a), each by clause. We have not located a sourced, current penalty or sanction figure for any of them that we can check against a primary source, so none is published.