{"entity":{"id":"idea-data-silent-trial-before-deployment","kind":"idea","name":"A mandatory silent (shadow) trial before any cancer AI goes live","aka":[],"tldr":"Before an AI tool is allowed to influence care at a hospital, it would run invisibly alongside clinicians for months so its real-world performance at that site is known first.","summary":"Shadow deployment (the model runs on live data but its outputs are hidden and compared with clinicians and outcomes) is standard practice at a few pioneering centres but not required. The proposal makes a pre-specified silent trial (minimum case numbers, pre-declared performance thresholds, subgroup analysis, comparison with local clinicians) a condition of go-live at each site, reported to the registry, as the AI equivalent of laboratory method verification before a new assay is used.","asOf":"2026-09-08","links":[{"label":"Bottleneck evidence (AI that is built but not validated or deployed): Wu et al., How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals (Nature Medicine 2021)","url":"https://doi.org/10.1038/s41591-021-01312-x"}],"tags":[],"related":["idea-data-drift-monitoring-standard"],"cancers":[],"sections":["ai-computation"],"technologies":[],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":[],"trials":[],"people":[],"bottlenecks":["b-ai-validation"],"keyPapers":["paper-wu-nat-med"],"journals":[],"dependsOn":[],"notes":[],"hypothesis":"Site-level silent trials will identify locally unacceptable performance in a meaningful share of deployments that passed regulatory clearance, preventing harm and building trust at sites where the tool passes.","rationale":"Clinical laboratories must verify every new assay locally before use, because performance depends on local conditions; AI has the same dependence and no equivalent rule.","test":"Require silent trials for all AI deployments in one hospital network for two years; report the proportion failing local thresholds and the reasons.","maturity":"early-clinical","actor":"clinic","cost":"small","horizonYears":1},"route":"/ideas/idea-data-silent-trial-before-deployment/","neighbours":{"idea":[{"id":"idea-data-drift-monitoring-standard","kind":"idea","name":"A standard for monitoring AI performance drift with pause thresholds","route":"/ideas/idea-data-drift-monitoring-standard/"}],"section":[{"id":"ai-computation","kind":"section","name":"AI & Computation","route":"/fronts/ai-computation/"}],"bottleneck":[{"id":"b-ai-validation","kind":"bottleneck","name":"AI that is built but not validated or deployed","route":"/bottlenecks/b-ai-validation/"}],"paper":[{"id":"paper-wu-nat-med","kind":"paper","name":"How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals","route":"/key-papers/paper-wu-nat-med/"}]}}