{"entity":{"id":"idea-data-neutral-ai-evaluator","kind":"idea","name":"A neutral public evaluator for cancer AI, on the model of NIST","aka":[],"tldr":"Create an independent public body whose job is to test cancer AI tools against each other on locked-away data and publish the scores, so hospitals can buy on evidence.","summary":"Hospitals cannot compare AI vendors; each presents its own validation. A publicly funded evaluator, running the sequestered benchmarks, publishing head-to-head results, subgroup performance and robustness tests, and updating as models change, would make procurement evidence-based and give regulators an independent data source. Models exist in NIST's testing programmes and the UK's AI evaluation initiatives.","asOf":"2026-09-08","links":[{"label":"Bottleneck evidence (AI that is built but not validated or deployed): Wu et al., How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals (Nature Medicine 2021)","url":"https://doi.org/10.1038/s41591-021-01312-x"}],"tags":[],"related":["idea-data-sequestered-prospective-benchmarks"],"cancers":[],"sections":["ai-computation"],"technologies":[],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":[],"trials":[],"people":[],"bottlenecks":["b-ai-validation"],"keyPapers":["paper-wu-nat-med"],"journals":[],"dependsOn":[],"notes":[],"hypothesis":"Publication of independent head-to-head results will shift procurement toward better-performing models and cause under-performing products to leave the market within three years.","rationale":"Independent testing works where buyers cannot verify claims themselves (cars, appliances, biometrics); cancer AI has exactly this information asymmetry.","test":"Fund the evaluator to test one task (mammography AI) across all vendors; survey procurement decisions in the following two years for reference to the results.","maturity":"speculative","actor":"policy","cost":"medium","horizonYears":3},"route":"/ideas/idea-data-neutral-ai-evaluator/","neighbours":{"idea":[{"id":"idea-data-sequestered-prospective-benchmarks","kind":"idea","name":"Sequestered, prospectively collected benchmark datasets that no one can train on","route":"/ideas/idea-data-sequestered-prospective-benchmarks/"}],"section":[{"id":"ai-computation","kind":"section","name":"AI & Computation","route":"/fronts/ai-computation/"}],"bottleneck":[{"id":"b-ai-validation","kind":"bottleneck","name":"AI that is built but not validated or deployed","route":"/bottlenecks/b-ai-validation/"}],"paper":[{"id":"paper-wu-nat-med","kind":"paper","name":"How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals","route":"/key-papers/paper-wu-nat-med/"}]}}