Our error rates
Every answer in a dossier carries the rate at which that kind of answer was wrong on past clearances. This page is the register behind those figures: what each means, where it comes from, on what population, and when it was measured. Where the registry does not state something, this page says so rather than guessing.
By answer
Regulatory pathway
0.03%wrong
how often the right pathway was the first answer
past clearances kept out of training (7,028)
- Era
- not stated in the registry
- Layer
- an evaluation run — the registry does not say whether live or batch
- Judged by
- FDA’s own decision on the record
- Also
- De Novo cases caught 100%
measured 2026-07-17
Product code
4.5%wrong
how often the committed product code was defensible
past clearances from after the training period (6,217)
- Era
- 2024–2025
- Layer
- the live system, re-run for this check
- Judged by
- FDA’s own label, or a code a rule finds co-valid with it
- Also
- exact match with FDA’s label 91.7%routes to the correct regulation 98.6%
measured 2026-08-04
Predicate
0.002%wrong
how often the committed predicate could carry an SE argument
past clearances kept out of training (45,922)
- Era
- not stated in the registry
- Layer
- the live system, re-run for this check
- Judged by
- the legal-eligibility rules (marketed, not recalled, not barred)
- Also
- live re-run on 500 of 500 set-aside cases
measured 2026-08-05
Equivalence table
1.9%wrong
how often the equivalence table passed FDA’s acceptance screen
past clearances from the eSTAR era (37,999)
- Era
- the eSTAR era (2013 onward)
- Layer
- not stated in the registry
- Judged by
- the RTA policy’s Appendix A checks
- Also
- verdicts self-consistent 95.4%
measured 2026-07-31
Equivalence narrative
0%wrong
how often the narrative stated only what the table supports
narratives assembled from the comparison table (100)
- Era
- not stated in the registry
- Layer
- not stated in the registry
- Judged by
- a fabrication check against the table
measured 2026-07-31
Testing evidence
56.8%wrong
how often a case that truly needed clinical evidence was flagged
past clearances from after the training period (4,477)
- Era
- 2023 onward
- Layer
- an evaluation run — the registry does not say whether live or batch
- Judged by
- not stated in the registry
measured 2026-08-02
Cybersecurity
0%wrong
how often a cyber package was required and we said so
a 400-device reference set graded by two independent methods (375)
- Era
- not stated in the registry
- Layer
- not stated in the registry
- Judged by
- a rule-level check no qualifying device can pass unflagged, plus the reference set
- Also
- precision 87% (over-flagging is deliberate)
measured 2026-08-02
Review time
21.1%wrong
how often the real review time fell inside the 80% range
past clearances from after the training period (173,964)
- Era
- not stated in the registry
- Layer
- not stated in the registry
- Judged by
- whether the real decision date fell inside the range
- Also
- median miss of a point guess 72.7 days — which is why a range is shown
measured 2026-08-02
Likely FDA questions
26.1%wrong
how often FDA’s question was among our top three
past clearances outside the training sample (5,833)
- Era
- not stated in the registry
- Layer
- an evaluation run — the registry does not say whether live or batch
- Judged by
- FDA’s own decision on the record
- Also
- a no-information baseline would score 73.4%
measured 2026-08-02
Risk & hazards
13.1%wrong
how often the listed hazards included the one that later recalled
past clearances from after the training period (13,460)
- Era
- 2026–2008
- Layer
- an evaluation run — the registry does not say whether live or batch
- Judged by
- not stated in the registry
- Also
- precision 25% (over-flagging is deliberate)
measured 2026-08-27
Refuse-to-Accept rules
0%wrong
how often each checklist rule was correctly applied
the 28 RTA rules, each audited trigger by trigger (28)
- Era
- not stated in the registry
- Layer
- not stated in the registry
- Judged by
- not stated in the registry
measured 2026-08-02
eSTAR fields
0%wrong
how often every eSTAR field held a legal value
filled eSTAR fields across four device types (108)
- Era
- not stated in the registry
- Layer
- not stated in the registry
- Judged by
- not stated in the registry
measured 2026-08-10
The figures on the landing page
- 45,922 of 45,922
- On every one of 45,922 past clearances re-run, the predicate the system proposed was legally eligible — cleared before the subject device, never recalled, not barred under 513(i)(2).
- Source
- Reliability registry R3 — legally eligible predicate (SE-VALID@1)
- Population
- 45,922 past clearances (batch, graded against bands A∪B) plus a live re-run of 500 held-out cases
- Method
- batch graded vs Band A∪B; live 500/500 held-out
- Measured
- 2026-08-05
- Re-measured
- re-measured on every registry release
- 95.5–100% on the eight answers that have a right answer
- Eight of the twelve answers have a single right answer, and the product gives it 95.5% to 100% of the time — five of the eight at 100%. The other four are not accuracies and are not stated as one: three are recall at a declared operating point (hazard symptoms; the clinical-evidence flag at 22.2% volume; FDA's question inside our top three) and one is a calibrated range (the review-time 80% interval covered 78.9%, which is on target — an interval covering 100% would be a worse interval). The figure for each answer is shown beside it, in its own terms.
- Source
- Reliability registry R1–R13
- Population
- per answer — see /error-rates
- Method
- the measured rate for each metric, stated in the direction the metric was measured — accuracy as accuracy, recall as recall, interval coverage against its target
- Measured
- 2026-08-27
- Re-measured
- re-measured on every registry release
- wrong 3–5% of the timethe worked example’s band; each dossier shows its own
- For a product-code answer shown at this confidence, the measured error on past clearances in the same confidence band.
- Source
- Calibration bands (Venn-Abers) and the isotonic meta-gate
- Population
- 1,528 past clearances, out-of-fold
- Method
- out-of-fold isotonic (cross_val_predict); the displayed p(wrong) is floored to the measured error of its band so it never undercuts the truth
- Measured
- 2026-08-11
- Re-measured
- re-measured with the calibration set
- 98.6%
- How often the committed product code led to the correct regulation on past clearances.
- Source
- Reliability registry R2 basis
- Population
- live out-of-time clearances 2024–25 (6,217)
- Method
- live OOT 2024-25; exact ~91.7%, defensible ~95.5%; routes-to-correct-regulation 98.58%
- Measured
- 2026-08-04
- Re-measured
- registry release
- 36% of the last 4,981 dossiersas of this snapshot; each dossier shows its own live trailing figure
- The share of dossiers the system sent to a person instead of deciding alone, over the trailing window. Counted as a miss on our side.
- Source
- ESCALATED events on the audit chain (serve_audit_event)
- Population
- all served dossiers in the window
- Method
- count of escalations ÷ dossiers served
- Measured
- rolling
- Re-measured
- live
- 42.4%
- How often the committed predicate was FDA’s exact historical pick — lower than eligibility because several equally valid predicates usually exist.
- Source
- Reliability registry R3 caveat
- Population
- 45,922 past clearances
- Method
- exact K-number match against the cleared record
- Measured
- 2026-08-05
- Re-measured
- registry release
Escalation — sending a dossier to a person instead of deciding — is counted as a miss on our side, never as caution, and its rate is published in every hand-off. The system card describes the models, data and limits behind these numbers: the system card.