{"entity":{"id":"idea-tr2-ai-external-validation-registry","kind":"idea","name":"A registry of external validation datasets for cancer AI models, with mandatory reporting","aka":[],"tldr":"Cancer AI models are usually tested on data from the same hospital they were built on. A registry of independent test datasets, and a rule that every model reports performance on at least one, would show which models really work.","summary":"Most published cancer AI models lack external validation, and performance drops sharply on data from other institutions. A curated registry of held-out datasets across modalities (pathology, radiology, genomics) hosted by neutral custodians, with a submission protocol that returns performance metrics without releasing the data, would make external validation routine. Journals and regulators would require a registry validation for any clinical claim.","asOf":"2026-09-08","links":[{"label":"Bottleneck evidence (Preclinical results do not reproduce): Errington et al., Investigating the replicability of preclinical cancer biology (eLife 2021)","url":"https://doi.org/10.7554/eLife.71601"}],"tags":[],"related":["idea-tr2-open-cdx-validation-sets","idea-multimodal-foundation-model"],"cancers":[],"sections":[],"technologies":["digital-pathology-ai","pathology-foundation-model"],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":[],"trials":[],"people":[],"bottlenecks":["b-reproducibility","b-ai-validation"],"keyPapers":[],"journals":[],"dependsOn":[],"notes":[],"hypothesis":"Models validated through the registry will show a median performance drop of at least ten percentage points from internal to external validation, and the requirement will improve the external performance of subsequently published models.","rationale":"Held-out evaluation servers (as in machine learning benchmarks) prevent overfitting to the test set; medicine has the datasets but not the shared infrastructure.","test":"Establish registry datasets for three tasks (HER2 scoring, lung nodule malignancy, ctDNA variant calling); validate 50 published models; report the distribution of performance changes.","maturity":"early-clinical","actor":"data","cost":"medium","horizonYears":2},"route":"/ideas/idea-tr2-ai-external-validation-registry/","neighbours":{"idea":[{"id":"idea-multimodal-foundation-model","kind":"idea","name":"Patient-level multimodal foundation models for treatment selection","route":"/ideas/idea-multimodal-foundation-model/"},{"id":"idea-tr2-open-cdx-validation-sets","kind":"idea","name":"Public gold-standard datasets for validating every cancer biomarker test","route":"/ideas/idea-tr2-open-cdx-validation-sets/"}],"technology":[{"id":"digital-pathology-ai","kind":"technology","name":"Digital pathology & AI","route":"/technologies/digital-pathology-ai/"},{"id":"pathology-foundation-model","kind":"technology","name":"Pathology & radiology foundation models","route":"/technologies/pathology-foundation-model/"}],"bottleneck":[{"id":"b-ai-validation","kind":"bottleneck","name":"AI that is built but not validated or deployed","route":"/bottlenecks/b-ai-validation/"},{"id":"b-reproducibility","kind":"bottleneck","name":"Preclinical results do not reproduce","route":"/bottlenecks/b-reproducibility/"}],"roadmap":[{"id":"ai-oncology-roadmap","kind":"roadmap","name":"AI in oncology roadmap: pattern readers → foundation models → agents in the workflow","route":"/roadmaps/ai-oncology-roadmap/"}],"term":[{"id":"external-validation","kind":"term","name":"External validation","route":"/terms/external-validation/"}]}}