Create an independent public body whose job is to test cancer AI tools against each other on locked-away data and publish the scores, so hospitals can buy on evidence.
Hospitals cannot compare AI vendors; each presents its own validation. A publicly funded evaluator, running the sequestered benchmarks, publishing head-to-head results, subgroup performance and robustness tests, and updating as models change, would make procurement evidence-based and give regulators an independent data source. Models exist in NIST's testing programmes and the UK's AI evaluation initiatives.
Shares Sequestered, prospectively collected benchmark datasets that no one can train on, How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals, AI that is built but not validated or deployed.
Shares Sequestered, prospectively collected benchmark datasets that no one can train on, How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals, AI that is built but not validated or deployed.
Shares How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals, AI that is built but not validated or deployed.
Shares How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals, AI that is built but not validated or deployed.
Shares How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals, AI that is built but not validated or deployed.
Shares How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals, AI that is built but not validated or deployed.
Shares How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals, AI that is built but not validated or deployed.
Shares How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals, AI that is built but not validated or deployed.