{"entity":{"id":"idea-data-sequestered-prospective-benchmarks","kind":"idea","name":"Sequestered, prospectively collected benchmark datasets that no one can train on","aka":[],"tldr":"Keep test datasets locked away and collect them going forward, so AI claims are checked on data the developers have never seen and could not have memorised.","summary":"Public benchmarks leak into training sets and go stale; retrospective validation flatters models. The proposal is a set of sequestered evaluation datasets for key cancer AI tasks (mammography, lung nodules, prostate biopsy, HER2 scoring, ctDNA calls), collected prospectively from multiple sites and countries, held by a neutral body, with evaluation only via submission of the model or an API, and results published. NIST's face recognition testing and the MICCAI challenge model are precedents.","asOf":"2026-09-08","links":[{"label":"NIST FRTE (face recognition evaluation)","url":"https://www.nist.gov/programs-projects/face-technology-evaluations-frtefate"}],"tags":[],"related":[],"cancers":[],"sections":["ai-computation"],"technologies":["radiology-ai-screening","digital-pathology-ai","ctdna"],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":[],"trials":[],"people":[],"bottlenecks":["b-ai-validation"],"keyPapers":[],"journals":[],"dependsOn":[],"notes":[],"hypothesis":"Performance on sequestered prospective data will be materially lower than published performance for most models, and public reporting will shift developers toward robust training and honest claims.","rationale":"In face recognition, NIST's sequestered testing became the de facto standard buyers rely on; in medical imaging, external test sets consistently reveal performance drops that published papers omit.","test":"Stand up two sequestered benchmarks; evaluate all willing vendors; publish results alongside their published claims; repeat annually to measure whether the gap narrows.","maturity":"early-clinical","actor":"research","cost":"medium","horizonYears":2},"route":"/ideas/idea-data-sequestered-prospective-benchmarks/","neighbours":{"section":[{"id":"ai-computation","kind":"section","name":"AI & Computation","route":"/fronts/ai-computation/"}],"technology":[{"id":"radiology-ai-screening","kind":"technology","name":"AI in radiology","route":"/technologies/radiology-ai-screening/"},{"id":"digital-pathology-ai","kind":"technology","name":"Digital pathology & AI","route":"/technologies/digital-pathology-ai/"}],"term":[{"id":"ctdna","kind":"term","name":"Circulating tumour DNA (ctDNA)","route":"/terms/ctdna/"}],"bottleneck":[{"id":"b-ai-validation","kind":"bottleneck","name":"AI that is built but not validated or deployed","route":"/bottlenecks/b-ai-validation/"}],"idea":[{"id":"idea-data-neutral-ai-evaluator","kind":"idea","name":"A neutral public evaluator for cancer AI, on the model of NIST","route":"/ideas/idea-data-neutral-ai-evaluator/"},{"id":"idea-data-multisite-validation-precondition","kind":"idea","name":"External validation at five or more sites in two countries before clearance","route":"/ideas/idea-data-multisite-validation-precondition/"},{"id":"idea-data-ai-red-team-programme","kind":"idea","name":"Red-team programmes that attack cancer AI before patients do","route":"/ideas/idea-data-ai-red-team-programme/"}]}}