Artificial intelligence in cancer started as software that flagged spots on a mammogram. It now designs molecules, reads slides better than any single pathologist for some tasks, and is beginning to match patients to trials and draft the tumour board summary; the question is which of it will be proven to help.
Three strands of AI are converging on oncology. In discovery, structure prediction (AlphaFold 3, Boltz, Chai) and generative chemistry have produced the first AI-designed candidates in trials, and perturbation-scale single-cell datasets are training models that try to predict what a drug will do to a cell. In diagnosis, foundation models trained on millions of slides and scans (Virchow, Prov-GigaPath, UNI, TITAN, CT-FM) underpin the first AI tests cleared to predict treatment benefit (ArteraAI Prostate 2025, ArteraAI Breast 2026) and the first randomised evidence that AI reading improves screening (MASAI). In the clinic, language models are entering trial matching, documentation and tumour-board support, with radiotherapy auto-contouring as the most mature deployed use.
The gap between the thousands of published models and the handful in clinical use is the defining feature of the field. Prospective, ideally randomised, evidence that an AI-guided decision improves an outcome exists for a few tools; a regulatory route for models that keep updating, payment codes for AI-derived biomarkers, and data that can be shared or federated across hospitals are all unsettled.
This roadmap covers the whole stack from molecule to clinic; the companion roadmaps go deeper on the AI-assisted clinic and on the virtual cell.
The first cleared cancer AI was computer-aided detection for mammography in 1998, which marked suspicious regions for the radiologist and, in large observational studies, did not improve accuracy. Rule-based decision support for treatment recommendations was tried and mostly abandoned. The lesson that survived: an algorithm has to be evaluated on the decision it changes, not on the pattern it finds.
Convolutional networks trained on labelled images matched specialists on narrow tasks. Paige Prostate (2021) became the first FDA-authorised AI for reading pathology slides; radiology triage tools for haemorrhage and embolism were cleared by the dozen; Sybil and Mirai predicted future lung and breast cancer from today's scan. Whole-slide scanning became routine in large centres, which made slide-level AI possible at all.
AlphaFold made protein structure a lookup rather than a two-year experiment; AlphaFold 3 (2024), Boltz and Chai extended it to drug-protein and antibody complexes, and RFdiffusion and ESM3 design proteins that never existed. Insilico's generative chemistry produced the first AI-discovered drug to reach phase 2, Isomorphic's first oncology candidate was cleared for trials, and Recursion and Xaira are betting that image and perturbation data can find targets no hypothesis would. None has yet produced an approved cancer drug, which is the honest benchmark.
Pathology models pretrained on millions of slides (Virchow, Prov-GigaPath, UNI and CONCH, H-optimus, TITAN) predict mutations, biomarkers and outcomes from a routine stain; MUSK adds clinical text. ArteraAI Prostate (2025) was the first AI test cleared to predict benefit from a treatment, and ArteraAI Breast followed in 2026. MASAI gave the first randomised evidence that AI-supported screening finds more cancers with less workload; Aidoc CARE (January 2026) was the first foundation-model triage platform cleared. Single-cell models (Geneformer, scGPT, State) and the Tahoe-100M dataset began the same arc for biology.
The first widely deployed AI in cancer care is not a diagnosis but a time-saver: auto-contouring of organs and tumours for radiotherapy planning now runs in hundreds of centres. Language models are being tested to match patients to trials from the record at the moment a treatment is chosen, to draft tumour-board summaries and pathology reports, and to answer patient questions under supervision. Federated learning lets models train across hospitals without moving data. The evidence standard for each is still being written.
The field has thousands of retrospective models and a handful of prospective trials. The infrastructure being proposed: a registry of external validation datasets with mandatory reporting, AI-first reading for high-volume common diagnoses with pathologists handling exceptions, every routine CT checked opportunistically for early cancer signs with a tracked pathway, AI central reads to cut trial endpoint cost, and digital twins as virtual control arms where a randomised control is unethical. Regulators are building predetermined change control plans so that models can update without re-clearance.
The two long-range bets are a multimodal model that reads slides, scans, genomics and the record to recommend and monitor treatment, and a virtual cell accurate enough to run a drug experiment in silico before it is run in a dish. Both depend on data at a scale no single institution holds, on validation standards that do not yet exist, and on liability and consent questions that are open today. The companion roadmaps on the AI clinic and the virtual cell follow each in detail.
The bottleneck is not model quality but the path from a published model to a deployed one: prospective evidence, external validation, regulatory status for updating models, payment, and data that can be shared. Records, scans and genomes sit in silos; real-world outcomes are weakly recorded, so there is little to learn from; and the workforce that would supervise AI is already short. Compute and model platforms are the one input that is not scarce.
Every era's records, trial outcomes and papers, and every watch item, as JSON.
Probability ranges are named estimates that the claim is borne out on roughly a five-year horizon. They are meant to be argued with: propose a revision with your name and reasoning via a pull request to src/data/confidence.ts.
Loading this step…
Loading this step…
Loading this step…
Loading this step…
Loading this step…
Loading this step…
Loading this step…
Loading this step…
Shares Prov-GigaPath (Microsoft, Providence), MUSK (Stanford, vision-language pathology), Virchow / Virchow2 (Paige, MSK), ArteraAI Breast.
Shares PathAI, Owkin, Paige AI, Patient-level multimodal foundation models for treatment selection.
Shares Massive Bio, Flatiron Health and Foundation Medicine Clinico-Genomic Database, Trial Library, Trial matching inside the electronic record at the moment a treatment is chosen.
Shares Arc Virtual Cell Atlas, Tahoe-100M, State (Arc Institute perturbation model), Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell.
Shares Owkin, Patient-level multimodal foundation models for treatment selection, Pathology & radiology foundation models, AI that is built but not validated or deployed.
Shares AI-assisted central imaging reads to cut endpoint cost and variability, AI-first reading for high-volume common cancer diagnoses, pathologist for the exceptions, Mammography & tomosynthesis, AI that is built but not validated or deployed.
Shares Owkin, Pathology & radiology foundation models, AI that is built but not validated or deployed, AI in radiology.