Thousands of cancer AI models are published; a handful are in clinical use, and fewer have shown they help patients.
Machine learning models for cancer detection, pathology, prognosis and treatment selection are published by the thousand, but almost all are evaluated retrospectively on data from the institution that built them. Of AI-enabled devices cleared by the FDA up to 2020, nearly all were evaluated only retrospectively and most on a single site, and among deep-learning studies comparing AI with clinicians, only a handful were prospective and two were randomised. Retrospective accuracy is not clinical benefit: models drift as scanners, populations and practice change, integration into workflow is costly, liability is unresolved, and reimbursement rarely exists. The MASAI trial of AI-supported mammography screening is one of the first randomised demonstrations that an AI tool can safely change a cancer pathway. Prospective and randomised evaluation, reporting standards, post-market monitoring for drift, and regulatory pathways for models that keep learning are the requirements for AI to move from papers into care.
Thousands of cancer AI tools have been tested on old data; almost none in a proper trial. Fund the trials, with endpoints that matter to patients.
Hospitals could train shared AI models on all their patients' scans and records without any data leaving the building, and jointly own the results, if someone built and governed the network.
Make clear who is responsible when an AI tool contributes to a mistake: protect doctors who use approved tools as intended, and hold makers responsible for the tool's performance.
Before an AI tool is allowed to influence care at a hospital, it would run invisibly alongside clinicians for months so its real-world performance at that site is known first.
Test the large language models doctors and patients are already using against a continually refreshed set of cancer questions, scoring not just correct answers but whether the sources they cite are real and support the claim.
Create an independent public body whose job is to test cancer AI tools against each other on locked-away data and publish the scores, so hospitals can buy on evidence.
Companies, hospitals and funders would form a consortium, like the Structural Genomics Consortium or IMI, to train one multimodal AI on scans, slides, genomes and outcomes from millions of patients by federated training across dozens of health systems, with the data never leaving the hospitals. Members would share the base model and compete on applications built on it.
A free web service where any app or hospital system can ask 'what is the recommended treatment for this exact situation today' and get a cited, versioned answer.
Patients now ask AI assistants about their cancer. Test those assistants regularly on real questions, publish the scores, and certify the ones that meet the bar.
Like a trial registry, every AI tool used on real patients would be listed publicly with what it is for, what data it was trained on, how well it performed and which version is running where.
AI tools that write clinic notes are spreading fast in cancer clinics. Test them properly: do they save time, do they make mistakes about drugs and doses, and do patients notice a difference?
Test head to head whether an AI that reads the record and the evidence recommends treatments as well as a panel of experts, and whether patients do as well.
Cancer AI models are usually tested on data from the same hospital they were built on. A registry of independent test datasets, and a rule that every model reports performance on at least one, would show which models really work.
Let AI tools that improve as they learn be used under close supervision in a few hospitals, with pre-agreed rules for what changes are allowed and how they are checked.
Regulators are starting to use AI to read dossiers faster. If they shared one tool, it could show where their questions overlap and where they truly disagree.
Pathology AI is cleared on uneven evidence, often without showing that pathologists using it do better than without. The proposed standard has two stages: a pre-registered, fully crossed multi-reader multi-case study comparing pathologist plus AI with pathologist alone, then a prospective deployment study measuring turnaround, tumour board discordance and treatment changes.
Set common rules for how hospitals check that an AI tool still works as the scanners, patients and practices around it change, and when it must be switched off.
Train a model on millions of experiments where genes and drugs were altered, so it can predict the effect of a new combination without running the experiment.
Most screening CT scans are normal. Letting a validated AI clear them, and sending only flagged scans to a radiologist, would let screening scale without more radiologists.
Most lung nodules on CT are harmless but trigger years of follow-up scans. A validated AI score could discharge low-risk nodules immediately.
Whether a lesion is called precancer or cancer varies between pathologists, and over time the bar has drifted lower. AI reference reads could hold the line.
Measuring tumours on scans for trials is slow, expensive and inconsistent between readers. Software that measures lesions and flags changes, checked by a radiologist, could make trial endpoints cheaper and more reliable.
MYC, fusion oncoproteins and transcription factors have shapeless, flexible regions that drugs cannot grip. Deep-learning protein design tools such as RFdiffusion may be able to invent binders that clamp them, for use as degradation handles, intrabodies or targeting domains for CAR and bispecific therapies rather than as drugs themselves.
Let validated AI make the first read on routine, high-volume samples like cervical smears and standard breast biopsy stains, so scarce pathologists spend their time on the difficult cases.
Hospitals buy multi-million-dollar surgical robots and AI tools with little proof they help patients. An independent body would run the comparative trials, and payers would only pay premiums for what is shown to work.
Build a shared, openly available AI model that has learned how cancer cells respond to genetic and drug perturbations, so any lab can predict what a new drug or combination might do.
Rather than giving the same dose until the cancer grows, measure tumour DNA in blood every few weeks and let a validated algorithm raise, lower, pause or switch drugs to keep the cancer suppressed for longer.
Cancer AI tools are approved on old test data and then never checked again. Require every deployed tool to report its real-world performance continuously, in public.
When a computer suggests a treatment, it should show the doctor the specific trial result and guideline sentence behind the suggestion, so it can be checked and trusted.
Autologous cell therapy batches fail more often than any other medicine because each patient's starting cells behave differently and the process runs without feedback. Inline sensors for metabolites, cell counts and cytokines, feeding models that adjust feeding and harvest timing in real time, could rescue batches that would otherwise be discarded.
Build a computer model of each patient's cancer that forecasts how it will respond to each treatment option, and prove it by writing the forecast down before the real result is known.
Scan the millions of cancer slides already sitting in hospital basements and connect each to what happened to the patient, creating the world's largest training set for pathology AI.
Many countries have one oncologist for millions of people. Train nurses and general doctors to deliver protocolised cancer care with software checks and remote specialist oversight.
Whenever an AI tool gives a result about a patient, the hospital system would permanently record what it saw, which version it was, what it said and what the doctor did with it.
No cancer AI would be approved until it has been tested on patients from at least five different hospitals in at least two countries, none of which contributed training data.
Train one AI on pathology slides and radiology scans from dozens of hospitals without any hospital sharing its images: the model travels to the data. Federated learning has worked for glioblastoma segmentation across 70-plus sites, yet almost every clinical model is still trained at one or two institutions, so a persistent shared training infrastructure is proposed.
Flu vaccines are chosen by predicting which virus strains will dominate next season. The same forecasting maths could predict which resistance mutation a patient's tumour will develop next.
Simulate trials of drug combinations in populations of virtual patients to decide which real trials to run, and keep score of how often the simulations were right.
Melanoma diagnoses have soared while deaths barely changed, a sign of overdiagnosis. AI skin apps should be judged on whether dangerous thick melanomas fall, not how many spots they flag.
Once an AI tool is in use, its maker and the hospital would have to report regularly how it is actually performing on real patients, and the reports would be public.
Every AI tool would have to report how well it works for women and men, different ethnic groups, ages, scanner types and hospitals, not just an overall score.
Thousands of papers extract 'radiomic' features from scans to predict outcomes, but the features change with scanner settings. Journals should require standard compliance before any clinical claim is made.
Build and certify a single free tool that strips names and identifying marks from cancer scans and pathology slides, so every hospital stops writing its own.
Much of a pathologist's day is preparation, measuring and describing specimens. Trained assistants can do that, and AI can pre-screen slides, so each pathologist reports far more cancers.
Every patient would be able to see which AI tools were used in their diagnosis or treatment plan, what they do, how well they work and how to question them.
Health systems would pay for AI tools that have shown in trials that they help patients, and pay nothing for tools that have not, giving makers a reason to run the trials.
Dozens of trials have collected immune, genomic and imaging data on the same drugs. Nobody can analyse them together, so the answer stays hidden in fragments.
A cancer blood test improves every year, but a ten-year trial tests the old version. Regulators and sponsors could agree in advance how updates are validated and carried into the result.
AI systems claim to find new uses for old drugs, but their predictions are rarely tested fairly. Publish their cancer predictions in advance and score them against trial results.
Anyone building a new test for HER2, PD-L1 or tumour DNA should be able to check it against the same public reference set. Today each developer validates on private data nobody can inspect.
If public or charity money paid to build a cancer AI model, the model itself (not just a paper about it) must be released so others can test, improve and use it.
Prostate MRI is good at telling you whether a man has a serious cancer and poor at telling you where all of it is. It missed at least one significant tumour in a third of men in the study that checked it against the whole removed prostate. Services that treat part of the gland, or follow men on imaging alone, are relying on the number they do not measure.
Every staging scan contains a precise measure of muscle mass that nobody looks at. Software could report it automatically and flag patients heading for wasting.
Pay independent experts to try to break cancer AI tools with unusual images, rare cases, bad scans and data shifts, and publish what breaks them.
AI for screening should be judged on whether it finds dangerous cancers earlier and misses fewer, not just on whether it agrees with radiologists on old images.
Just as drugs are withdrawn when they prove unsafe, AI tools should have clear triggers for being switched off, and someone responsible for pulling the switch.
No one keeps score of which laboratory models actually predicted what happened in patients. A public scoreboard would show which models to trust.
Keep test datasets locked away and collect them going forward, so AI claims are checked on data the developers have never seen and could not have memorised.
Whether immune cells are next to cancer cells matters more than how many there are. Turning that spatial picture into a reliable, standardised test would predict response better.
Drawing targets and planning radiotherapy takes hours of scarce expert time. Properly tested AI could do much of it, letting the same staff treat far more patients, if regulators and payers set clear rules for proving and paying for it.
AI is starting to decide which patients get which cancer drug. Every change to the software should be tested against a fixed public set of cases before it is used on patients.
Build a computer model of each patient's cancer and body that simulates how different treatments would go, and prove in a proper trial that choosing treatment with the model helps.
AI can take over one reader's work in double-reading screening programmes while finding more cancers. Whether the extra cancers found are ones that would have harmed women, and whether interval cancers fall, is the question the trial's primary endpoint will answer.
The shape of nearly every protein is now available to any researcher in seconds instead of years, which shortens the path from a cancer target to a designed molecule. It does not by itself produce drugs: binding pockets, dynamics and cellular context still need experiment.
One bottleneck page and 22 idea pages on OnCo cite this paper by its DOI; this record gives the citation a page of its own so a reader can follow it without leaving OnCo. Read the abstract above alongside the citing pages listed under Related; the record was created automatically from the Europe PMC entry and its figures have not been checked by hand.
One bottleneck page on OnCo cites this paper by its DOI; this record gives the citation a page of its own so a reader can follow it without leaving OnCo. Read the abstract above alongside the citing page listed under Related; the record was created automatically from the Europe PMC entry and its figures have not been checked by hand.
Shares A randomised trial of AI scribes in oncology clinics measuring errors and time, Double oncology capacity in low-resource settings with task-shifting and AI decision support, Pathologist assistants plus AI triage to multiply pathologist capacity, A randomised trial of AI-generated treatment recommendations versus tumour boards.
Shares ArteraAI Breast, ArteraAI Prostate, Paige AI, Patient-level multimodal foundation models for treatment selection.
Shares Every AI output logged in the record with input hash, version and clinician response, Digitise the nation's pathology slides and link them to outcomes, One certified open-source de-identification pipeline for scans and slides, Federated training of pathology and radiology models across hospitals.
Shares Judge skin cancer AI by the thick melanomas it prevents, not the thin ones it finds, AI malignancy scores to end repeat scans and biopsies for benign lung nodules, AI second reads to stop borderline lesions being upgraded to cancer, Require stage-shift or interval-cancer endpoints for AI in cancer screening.