{"entity":{"id":"idea-data-ai-red-team-programme","kind":"idea","name":"Red-team programmes that attack cancer AI before patients do","aka":[],"tldr":"Pay independent experts to try to break cancer AI tools with unusual images, rare cases, bad scans and data shifts, and publish what breaks them.","summary":"Robustness of medical AI to artefacts, rare presentations, adversarial inputs and distribution shift is poorly characterised. The proposal funds standing red teams (imaging physicists, pathologists, security researchers) that stress-test cleared and pre-clearance cancer AI with curated adversarial and edge-case corpora, publish failure modes in a common taxonomy, and feed results to the registry and developers, as is done for cybersecurity and increasingly for general-purpose AI.","asOf":"2026-09-08","links":[{"label":"Bottleneck evidence (AI that is built but not validated or deployed): Wu et al., How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals (Nature Medicine 2021)","url":"https://doi.org/10.1038/s41591-021-01312-x"}],"tags":[],"related":["idea-data-sequestered-prospective-benchmarks"],"cancers":[],"sections":["ai-computation"],"technologies":[],"targets":[],"drugs":[],"companies":[],"institutions":[],"pathways":[],"terms":[],"trials":[],"people":[],"bottlenecks":["b-ai-validation"],"keyPapers":["paper-wu-nat-med"],"journals":[],"dependsOn":[],"notes":[],"hypothesis":"Red-teaming will uncover clinically relevant failure modes in most tested models that were not disclosed in their validation, and disclosure will lead to fixes or labelling changes.","rationale":"Every mature safety-critical field uses adversarial testing; medical AI relies on developers' own validation, which is structurally blind to what the developers did not think of.","test":"Red-team ten cancer AI tools over one year; publish findings; track developer responses and label changes within a further year.","maturity":"speculative","actor":"research","cost":"small","horizonYears":1},"route":"/ideas/idea-data-ai-red-team-programme/","neighbours":{"idea":[{"id":"idea-data-sequestered-prospective-benchmarks","kind":"idea","name":"Sequestered, prospectively collected benchmark datasets that no one can train on","route":"/ideas/idea-data-sequestered-prospective-benchmarks/"}],"section":[{"id":"ai-computation","kind":"section","name":"AI & Computation","route":"/fronts/ai-computation/"}],"bottleneck":[{"id":"b-ai-validation","kind":"bottleneck","name":"AI that is built but not validated or deployed","route":"/bottlenecks/b-ai-validation/"}],"paper":[{"id":"paper-wu-nat-med","kind":"paper","name":"How medical AI devices are evaluated: limitations and recommendations from an analysis of FDA approvals","route":"/key-papers/paper-wu-nat-med/"}]}}