Keep test datasets locked away and collect them going forward, so AI claims are checked on data the developers have never seen and could not have memorised.
Public benchmarks leak into training sets and go stale; retrospective validation flatters models. The proposal is a set of sequestered evaluation datasets for key cancer AI tasks (mammography, lung nodules, prostate biopsy, HER2 scoring, ctDNA calls), collected prospectively from multiple sites and countries, held by a neutral body, with evaluation only via submission of the model or an API, and results published. NIST's face recognition testing and the MICCAI challenge model are precedents.
Shares AI that is built but not validated or deployed, AI in radiology, Digital pathology & AI.
Shares AI that is built but not validated or deployed, AI in radiology, Digital pathology & AI.
Shares AI that is built but not validated or deployed, AI in radiology, Digital pathology & AI.
Shares A neutral public evaluator for cancer AI, on the model of NIST, Red-team programmes that attack cancer AI before patients do, External validation at five or more sites in two countries before clearance, AI that is built but not validated or deployed.
Shares AI that is built but not validated or deployed, AI in radiology, Digital pathology & AI.
Shares AI that is built but not validated or deployed, AI in radiology.
Shares AI that is built but not validated or deployed, Digital pathology & AI.
Shares AI that is built but not validated or deployed, Digital pathology & AI.