Loading
44 foundation and risk models and 21 datasets from the corpus, with the fields that let you compare them: parameters, modality, training data, whether the weights can be downloaded (27 are open), licence, a reported benchmark and the paper. Figures are the developers' own; blanks are unverified, not zero.
| Training data / size | Reported result / consent | Paper / weights | Trained on / used by | |||||
|---|---|---|---|---|---|---|---|---|
Aidoc CARE (clinical radiology foundation model) Aidoc CARE is one radiology foundation model, pretrained on CT scans without labels, whose task-specific heads have each been FDA-cleared to flag urgent findings in emergency scans so radiologists read those first. Its oncology relevance is indirect, catching incidental masses; the regulatory evidence covers triage, not diagnostic accuracy for tumours. | none | Not disclosed; underpins FDA-cleared triage products. | none | none | none | |||
AlphaGenome Reads a million letters of DNA at once and predicts how a mutation changes gene regulation, splicing and chromatin. | none | Human and mouse reference genomes with thousands of functional genomics tracks; 1 megabase input at single-base resolution. | Matched or beat specialist models on 22 of 24 sequence prediction and 24 of 26 variant-effect tasks (DeepMind). | paper, weights | none | |||
Arc Virtual Cell Atlas Arc's growing library of cell data, the fuel for virtual cell models. | none | Hundreds of millions of cells (observational and perturbational) | Harmonised public datasets; per-dataset licences apply. | portal | State | |||
Atlas (Aignostics, Mayo Clinic, Charité) Atlas is a pathology foundation model trained on 1.2 million slides from two of the world's largest hospitals. | none | 1.2 million slides from Mayo Clinic and Charité across scanners and stains. | Top scores on public benchmark tasks reported in the arXiv paper. | arXiv | none | |||
BioEmu (Microsoft) BioEmu is a Microsoft generative diffusion model that predicts the range of shapes a protein moves between, not one static structure, thousands of times faster than molecular dynamics simulation. For cancer drug discovery that can reveal transient pockets, as in KRAS, that static predictors miss, but its outputs are approximate and validated mainly on small proteins. | none | Molecular dynamics ensembles and experimental folding free energies. | none | paper, weights | none | |||
Cell2Sentence / C2S-Scale (Yale, Google) Turns a cell's gene expression into a sentence so a normal language model can reason about it; a 27-billion-parameter version proposed a cancer immunotherapy idea that was confirmed in the lab. | 27 B | Cell sentences from more than 50 million cells plus biological text, on Gemma-2 backbones up to 27B. | Predicted that silmitasertib raises antigen presentation under low interferon; validated in vitro (paper). | bioRxiv, weights | CZ CELLxGENE / Human Cell Atlas | |||
CT-FM (whole-body CT foundation model) A model pretrained on 148,000 CT scans to segment organs and triage findings. | none | 148,000 whole-body CT scans, self-supervised. | none | arXiv, weights | The Cancer Imaging Archive, NCI Imaging Data Commons | |||
ESM3 (EvolutionaryScale) ESM3 is a generative protein model that designed a working fluorescent protein far from any natural sequence. | 98 B | 2.78 billion protein sequences with structure and function tokens. | Generated esmGFP, a functional fluorescent protein 58% identical to its nearest known relative (Science 2025). | paper, weights | none | |||
Evo 2 (Arc Institute, NVIDIA) A DNA language model trained on 9.3 trillion bases that can flag cancer-causing BRCA1 variants without being told about them. | 40 B | 9.3 trillion nucleotides across all domains of life (OpenGenome2); 7B and 40B parameter models, 1 million token context. | Zero-shot BRCA1 variant pathogenicity prediction competitive with supervised methods (paper). | bioRxiv, weights | none | |||
Midnight (kaiko.ai) Midnight is a pathology model that matched the leaders while training on far fewer slides. | none | 12,000 TCGA slides (Midnight-12k) with a combined DINOv2 and high-resolution objective. | Leading scores on the eva pathology benchmark suite with a fraction of competitors' training slides (kaiko.ai). | weights or model card | TCGA / NCI Genomic Data Commons, Pathology AI benchmarks | |||
MUSK (Stanford, vision-language pathology) A model that reads slides and clinical text together to predict who will respond to immunotherapy. | none | 50 million pathology images and 1 billion pathology-related text tokens. | Better prediction of immunotherapy response and prognosis than single-modality models across cancers (paper). | paper, weights | none | |||
State (Arc Institute perturbation model) Predicts how cells will respond to a drug or gene knockout, trained on over 100 million perturbed cells. | none | State transition model: over 100 million perturbed cells including Tahoe-100M; embedding model: 167 million human cells. | Reference entry for the Arc Virtual Cell Challenge (2025). | bioRxiv, weights | Tahoe-100M, Arc Virtual Cell Atlas | |||
Tahoe-100M Tahoe-100M is the biggest single-cell dataset ever released, built to teach AI how cancer cells respond to drugs. | none | 100 million single-cell profiles; 1,100 drugs; 50 cell lines | Cell lines only; CC BY release. | portal | State | |||
TranscriptFormer and rBio (CZI virtual cell models) CZI's open cross-species cell models and a reasoning model trained on them. | none | 112 million cells across 12 species. | none | weights or model card | CZ CELLxGENE / Human Cell Atlas | |||
AlphaFold 3 Predicts the 3D shape of proteins together with DNA, RNA, small molecules and antibodies, the starting point for much modern drug design. | none | Protein Data Bank structures and distillation sets across proteins, nucleic acids, ligands and ions. | At least 50% better ligand-interaction accuracy than prior methods on PoseBusters (paper). | paper, weights | none | |||
Boltz-1 / Boltz-2 (MIT, open) Open-source structure models that match AlphaFold 3, with Boltz-2 also predicting how strongly a drug binds. | none | PDB and distillation data; Boltz-2 adds binding-affinity training. | Boltz-1 matched AlphaFold 3 accuracy; Boltz-2 affinity prediction approached FEP+ at about 1,000 times lower cost (preprint). | bioRxiv, weights | none | |||
CellFM CellFM is an 800-million-parameter single-cell model trained on 100 million human cells. | 800 M | About 100 million human cells. | none | bioRxiv, weights | none | |||
Chai-1 / Chai-2 Structure and antibody-design models from Chai Discovery, with Chai-2 reporting high zero-shot antibody hit rates. | none | PDB-derived structures; Chai-2 trained for antibody and binder design. | Chai-2: about 16% zero-shot binder hit rate across dozens of targets in wet-lab tests (company report). | bioRxiv, weights | none | |||
CHIEF (Harvard, Yu Lab) A pathology model trained across 19 cancer types that predicts survival and mutations from slides. | none | Pretrained on 15 million tiles, then 60,530 slides; validated on 19,400 slides from 24 hospitals. | Cancer detection accuracy up to 94% and improvement of about 36% over prior deep-learning methods on external cohorts (paper). | paper, weights | TCGA / NCI Genomic Data Commons, CPTAC | |||
Foresight (generative EHR model) A model trained on millions of hospital records that forecasts a patient's next diagnoses. | none | Coded EHR timelines from King's College Hospital and South London and Maudsley; Foresight 2 extended to more than 5 million patients. | none | paper | none | |||
H-optimus (Bioptimus) An open 1.1-billion-parameter pathology model from a French startup, among the strongest on public benchmarks. | 1.1 B | Hundreds of millions of tiles from 500,000 slides (H-optimus-0). | none | weights or model card | none | |||
Hibou (HistAI) Hibou is a family of open pathology foundation models under a permissive licence. | 307 M | Over 1 million slides (Hibou-B 86M parameters; Hibou-L 307M). | none | arXiv, weights | none | |||
Med-Gemini and MedLM (Google) Google's medical versions of its Gemini models, able to reason over text, images, and long records. | none | Gemini models fine-tuned on medical text, images and genomics. | 91.1% on MedQA (USMLE) in the arXiv report. | arXiv | none | |||
MedSAM / SAM-Med3D (segment anything for medicine) Adaptations of Meta's Segment Anything model that outline tumours and organs on any scan with a click. | none | 1.57 million image-mask pairs across 10 imaging modalities and more than 30 cancer types. | Median Dice above specialist models on internal and external validation (paper). | paper, weights | The Cancer Imaging Archive | |||
Merlin (Stanford abdominal CT vision-language model) Merlin is a model trained on 15,000 CT scans with their reports that can find and describe hundreds of findings. | none | 6 million images from 15,331 abdominal CTs with 6 million EHR diagnosis codes and 1.8 million reports. | none | arXiv, weights | none | |||
Nicheformer (spatial single-cell) Nicheformer is a model trained on both dissociated and spatial data so it learns how a cell's neighbourhood shapes it. | none | 110 million cells: 57 million dissociated and 53 million spatially resolved. | none | bioRxiv, weights | none | |||
Nucleotide Transformer (InstaDeep) DNA language models trained on thousands of genomes for variant and regulatory prediction. | 2.5 B | 3,202 human genomes and 850 species genomes. | none | paper, weights | none | |||
Phenom-2 and Recursion OS A model trained on billions of cell microscopy images to read what a drug or gene knockout does to a cell. | 1.9 B | 8 billion cell-painting images (Recursion). | none | none | none | |||
Phikon / Phikon-v2 (Owkin) Owkin's open pathology models trained on TCGA and its federated hospital network. | 307 M | Phikon: 43 million tiles from TCGA. Phikon-v2: 460 million tiles from 58 million slides of 30 cancer types (PANCAN-XL). | none | arXiv, weights | TCGA / NCI Genomic Data Commons | |||
PLUTO (PathAI) PLUTO is PathAI's compact pathology foundation model, a vision transformer pretrained at several magnifications on 195 million tiles from 158,000 slides, so one network serves slide-level and biomarker quantification tasks at whatever resolution each needs. It runs inside PathAI's AISight product, but its weights are proprietary, so outside groups cannot benchmark or adapt it. | none | 195 million tiles from 158,000 slides across more than 50 sites, multi-scale. | Powers PathAI AISight and biomarker quantification products; results in the arXiv report. | arXiv | none | |||
Prov-GigaPath (Microsoft, Providence) An open pathology model trained on 1.3 billion image tiles from a US health system, modelling whole slides at gigapixel scale. | 1.3 B | 1.3 billion tiles from 171,189 whole slides from more than 30,000 Providence patients. | Best on 25 of 26 tasks in the paper, including mutation prediction and cancer subtyping. | paper, weights | TCGA / NCI Genomic Data Commons | |||
scFoundation (BioMap) scFoundation is a 100-million-parameter model trained on 50 million cells, from China's BioMap. | 100 M | Over 50 million cells across all about 19,000 human genes. | none | paper, weights | none | |||
scGPT A GPT-style model for single-cell data that predicts cell types, perturbation responses, and gene networks. | none | 33 million cells (CELLxGENE). | none | paper, weights | CZ CELLxGENE / Human Cell Atlas | |||
Tempus multimodal models Models trained on Tempus's paired genomic, pathology, imaging and outcome data to predict response and prognosis. | none | Tempus clinico-genomic and imaging data; details in company publications. | none | none | none | |||
TITAN (whole-slide multimodal model) TITAN is a model that summarises a whole slide, not just tiles, and can write a draft pathology report. | none | 335,645 whole slides (Mass-340K) with vision-only and vision-language pretraining; slide-level embeddings. | Slide-level retrieval, prognosis and report generation; outperformed patch-based aggregation in the paper. | arXiv, weights | none | |||
UNI and CONCH (Harvard, Mahmood Lab) Two open academic pathology models: UNI reads tissue images, CONCH links images with pathology text. | 307 M | UNI: 100 million tiles from 100,000 slides across 20 tissue types (Mass-100K). CONCH: 1.17 million image-caption pairs. | UNI: best or tied-best on 34 clinical tasks in the paper; CONCH: zero-shot classification and retrieval across 14 tasks. | paper, weights | TCGA / NCI Genomic Data Commons, Pathology AI benchmarks | |||
Virchow / Virchow2 (Paige, MSK) A pathology foundation model trained on millions of slides that can detect cancer and predict biomarkers from an ordinary H&E slide. | 1.9 B | Virchow: 1.5 million slides from MSK (632M parameters). Virchow2 and Virchow2G: 3.1 million slides from MSK and global sites at mixed magnification. | Pan-cancer detection AUC 0.95 across 17 cancer types, including rare cancers, in the Virchow paper. | paper, weights | none | |||
AlphaMissense Scored all 71 million possible single-letter protein changes in humans as likely harmful or benign. | none | Fine-tuned from AlphaFold on population frequency data; scored all 71 million possible human missense variants. | Classified 89% of missense variants as likely benign or likely pathogenic (Science 2023). | paper, weights | none | |||
Geneformer Geneformer is a transformer trained on about 30 million single cells that encodes each cell as a ranked list of its genes, so deleting a gene in silico shows which genes matter in a disease. It was the first single-cell foundation model in general use, though benchmarks find only modest gains over linear baselines on some tasks. | none | Genecorpus-30M (about 30 million cells), later 95 million cells. | none | paper, weights | CZ CELLxGENE / Human Cell Atlas | |||
GenePT Uses text embeddings of gene descriptions from a general LLM to represent cells, and performs surprisingly well. | none | GPT-3.5 embeddings of NCBI gene summaries; no single-cell pretraining. | none | bioRxiv, weights | none | |||
RadFM (generalist radiology foundation model) An open generalist model that answers questions about 2D and 3D scans. | 14 B | MedMD: 16 million 2D and 3D scans with text. | none | arXiv, weights | none | |||
RFdiffusion / RFdiffusion2 and ProteinMPNN (Baker Lab) The tools that design entirely new proteins to bind a chosen target, now used for cancer binders and antibodies. | none | PDB structures; fine-tuned from RoseTTAFold. | Experimentally validated binders, symmetric assemblies and enzyme active sites across the paper's design tasks. | paper, weights | none | |||
Sybil (MIT/MGH lung cancer risk from CT) Predicts a person's six-year lung cancer risk from one low-dose CT, even when no nodule is visible. | none | Low-dose CT scans from the National Lung Screening Trial; validated at MGH and in Taiwan. | One-year lung cancer risk AUC 0.86 to 0.94 across validation sets (JCO 2023). | paper, weights | NCI CDAS: NLST and PLCO screening trial data | |||
Universal Cell Embedding (UCE) Universal Cell Embedding maps any cell from any species into one shared space without retraining. | none | 36 million cells across 8 species, using protein embeddings of genes. | none | bioRxiv, weights | CZ CELLxGENE / Human Cell Atlas | |||
Enformer and Borzoi (DeepMind, Calico) Models that predict how DNA sequence controls gene activity, used to interpret non-coding cancer mutations. | none | Enformer: 200 kb input, thousands of epigenomic tracks. Borzoi: 524 kb input, RNA-seq coverage across human and mouse. | none | paper, weights | none | |||
Mirai (MIT breast cancer risk from mammograms) Reads a mammogram to estimate five-year breast cancer risk, consistently across races and devices. | none | Mammograms from MGH; validated across seven hospitals in several countries. | Five-year breast cancer risk C-index 0.76 to 0.81 across sites (Sci Transl Med 2021). | paper, weights | none | |||
Patient digital twins A digital twin is a computer model of one patient's tumour and body, updated with each scan and blood test, used to forecast how the disease will respond to each option before it is tried. | none | A patient-specific state model is calibrated and continually updated from that patient's data, then run forward under alternative treatments to compare predicted outcomes. | none | Individual forecasts rather than averages | source | none | ||
All of Us Research Program America's answer to UK Biobank, built for diversity. | none | More than 250,000 whole genomes; target one million participants | Registered and controlled tiers via the Researcher Workbench; emphasis on under-represented groups. | portal | none | |||
AACR Project GENIE AACR Project GENIE is real-world tumour sequencing data shared by leading cancer centres. | none | More than 200,000 sequenced tumours from 19 institutions | De-identified clinico-genomic data via cBioPortal and Synapse; Biopharma Collaborative adds outcomes. | portal | none | |||
DepMap (Cancer Dependency Map) Which genes each cancer cell line cannot live without. The map of synthetic-lethal targets. | none | Genome-wide CRISPR and drug screens on more than 1,000 cell lines | Cell lines only; CC BY 4.0. | portal | none | |||
Evolutionary dynamics of drug resistance Mathematics from population genetics shows that resistant cells almost always exist before treatment in large tumours, and that combining drugs with different resistance mutations from the start can succeed where the same drugs in sequence fail. | none | Resistance arises as a stochastic branching process; the probability of pre-existing resistance to k drugs falls steeply with k when their resistance mutations are distinct, favouring simultaneous combinations. | none | Quantitative case for upfront combinations | source | none | ||
Solid stress and tumour mechanobiology models Growing tumours compress themselves and their surroundings; models of this solid stress explain collapsed vessels, poor drug delivery and stiff stroma, and point to drugs that soften the tumour so treatment can get in. | none | Tumour growth against confinement generates solid stress; continuum mechanics relates stress to vessel collapse and interstitial pressure, and matrix or cell depletion relieves it. | none | Explains drug-delivery failure in dense tumours | source | none | ||
Evolutionary game theory in cancer Cancer cells are treated as players whose success depends on what neighbouring cells do, which lets researchers predict how a tumour's mix of cell types shifts under treatment and design schedules that steer it. | none | Replicator dynamics: each cell type's frequency changes in proportion to its fitness relative to the population mean, with fitness set by a payoff matrix of pairwise interactions that treatment can alter. | none | Captures cell-cell interactions other models ignore | source | none | ||
Quantitative systems pharmacology (QSP) Mechanistic computer models that join a drug's pharmacology to the biology of the tumour and the body, used by developers and regulators to pick doses, predict combinations and explain why a trial failed. | none | Systems of differential equations link drug concentration to receptor occupancy, pathway activity, cell kill and clinical markers, calibrated to preclinical and clinical data and used to simulate untested regimens. | none | Accepted by regulators for dose justification | FDA Project Optimus | none | ||
The Cancer Imaging Archive (TCIA) The public archive of cancer scans that most radiology AI is trained and tested on. | none | More than 200 imaging collections (CT, MRI, PET, pathology) | De-identified; mostly CC BY, some collections restricted. | portal | MedSAM / SAM-Med3D, CT-FM | |||
Adaptive therapy (evolution-based dosing) Instead of hitting a tumour as hard as possible, adaptive therapy gives just enough drug to keep it in check and stops when it shrinks, so drug-sensitive cells survive to compete with resistant ones. A pilot trial in prostate cancer roughly doubled the time to progression on abiraterone. | none | Lotka-Volterra competition between sensitive and resistant clones: keeping a sensitive population alive imposes a fitness cost on resistant cells, delaying their takeover; dosing is adjusted to hold the tumour at a stable burden. | none | Doubled time to progression in the pilot trial with half the drug | source | none | ||
TCGA / NCI Genomic Data Commons The reference atlas of cancer genomes that most cancer biology since 2008 is built on. | none | More than 11,000 tumours across 33 cancer types (TCGA) plus TARGET and CPTAC | Open for processed data; dbGaP authorisation for raw sequence and germline. | portal | UNI and CONCH, Prov-GigaPath, CHIEF | |||
UK Biobank UK Biobank is the richest population cohort for linking genes, blood and imaging to who later develops cancer. | none | 500,000 participants; whole genomes, imaging, proteomics, linked cancer registry | Approved researchers only; broad consent with linkage; no return of individual results. | portal | none | |||
Agent-based and multicellular simulations Instead of equations for average behaviour, agent-based models simulate every cell as an individual with rules for dividing, moving, dying and signalling, producing virtual tumours in which immune attack, drug delivery and evolution can be watched and tested. | none | Discrete cells with rule-based behaviour interact with each other and with continuous fields of nutrients and drugs; population behaviour emerges rather than being assumed. | none | Captures spatial and cell-level heterogeneity | source | none | ||
Residual disease kinetics (BCR-ABL halving and ctDNA slopes) The speed at which a molecular marker falls during treatment predicts outcome better than a single level: BCR-ABL halving time in chronic myeloid leukaemia and circulating tumour DNA slopes in solid tumours are now used to judge response within weeks. | none | Exponential (often biphasic) decline of a tumour-derived marker whose rate constants reflect cell kill in different compartments; early slope predicts depth and durability of response. | none | Predicts outcome weeks into treatment | source | none | ||
Proliferation-invasion (reaction-diffusion) models of glioma Gliomas grow by both dividing and migrating through the brain, and a two-parameter equation fitted to a patient's MRI scans estimates how far invisible cells have spread, which can guide how much brain to irradiate and how fast the tumour will grow. | none | Partial differential equation dc/dt = grad(D grad c) + rho·c(1 - c/K): net growth from proliferation and spatial spread from diffusion, with D higher in white matter. | none | Patient-specific from routine MRI | source | none | ||
Angiogenesis and vascular normalisation models Models of how tumours recruit blood vessels, and of Rakesh Jain's idea that anti-angiogenic drugs at the right dose normalise rather than destroy those vessels, improving drug and oxygen delivery for a window of days. | none | Vessel growth follows chemotactic and haptotactic gradients of angiogenic factors; anti-angiogenic therapy at moderate dose reduces vessel density and leakiness, lowering interstitial pressure and improving perfusion transiently. | none | Explains combination benefit of anti-VEGF drugs | Jain 2005, Science | none | ||
Tumour-immune dynamics models Predator-prey style equations describe how immune cells hunt tumour cells, and they reproduce dormancy, escape and the delayed, sometimes explosive, responses seen with immunotherapy; they now help design combination and scheduling trials. | none | Coupled ordinary differential equations for tumour and immune populations with recruitment, killing, exhaustion and suppression terms, whose equilibria correspond to dormancy, escape or elimination. | none | Explains non-linear immunotherapy responses | source | none | ||
Tumour control and normal tissue complication probability (TCP and NTCP) Curves that turn a radiation dose into a probability: how likely the tumour is to be eradicated and how likely a nearby organ is to be damaged. They underlie dose constraints, dose escalation trials and comparisons between treatment plans. | none | TCP = exp(-N·S(D)) for N clonogens with survival S(D); NTCP as a sigmoid of the equivalent uniform dose to an organ with volume-effect parameters, fitted to clinical outcome data. | none | Basis of organ dose constraints | source | none | ||
Goldie-Coldman model of resistance Resistant cells arise by chance mutation as a tumour grows, so the chance of a cancer already containing resistant cells rises with its size. The model argued for treating early and for alternating non-cross-resistant drugs. | none | Probability that no resistant cell exists in a tumour of N cells with mutation rate u is about exp(-u·N); resistance is therefore expected once tumours exceed roughly 1/u cells. | none | First quantitative theory of acquired resistance | source | none | ||
Norton-Simon hypothesis and dose-dense chemotherapy Because tumours regrow fastest when small, the best way to finish them is to give the same chemotherapy doses closer together. The idea, from Larry Norton and Richard Simon, was proved in breast cancer by the CALGB 9741 trial and made two-weekly chemotherapy a standard. | none | Kill rate is proportional to the Gompertzian growth rate of the tumour, so regrowth between cycles is the enemy and shorter intervals matter more than higher doses. | none | Confirmed in randomised trials and meta-analysis | CALGB 9741, Citron 2003 | none | ||
Clonal evolution and branching models Peter Nowell's 1976 idea that a tumour is an evolving population of competing clones is now measured directly by sequencing several regions or repeated blood samples, and models of branching evolution predict which clones will drive relapse. | none | A tumour is a population under mutation, drift and selection; phylogenetic and branching-process models reconstruct its history from the variant allele frequencies observed across regions and time. | none | Directly supported by multi-region and serial sequencing | Nowell 1976, Science | none | ||
The four Rs and accelerated repopulation Radiotherapy works through repair, redistribution, reoxygenation and repopulation between fractions; the discovery that tumours speed up their regrowth during a course explained why gaps in treatment cost cures and led to accelerated schedules. | none | Tumour control depends on total dose, dose per fraction and overall time; loss of control per day of prolongation beyond a kick-off time reflects accelerated repopulation of clonogens. | none | Explains the cost of treatment gaps | source | none | ||
Pharmacokinetic and pharmacodynamic modelling Equations that describe how a drug's concentration rises and falls in the body and how that concentration translates into effect and toxicity; the reason doses are given per square metre, why some drugs are infused over days, and how children's doses are set. | none | Drug concentration follows first-order transfer between compartments; effect is a saturable (Emax) function of concentration; population variability is modelled as random effects around typical parameters. | none | Basis of every dosing regimen | Wikipedia | none | ||
Gompertzian tumour growth Tumours do not grow exponentially forever: growth slows as they enlarge, following a curve Benjamin Gompertz devised for human mortality in 1825 and Anna Kane Laird fitted to tumours in 1964. It explains why small tumours are the most chemosensitive and why doubling times lengthen. | none | dV/dt = a·V·ln(K/V): the growth rate is proportional to size times the logarithm of the distance from the carrying capacity K, giving a sigmoid curve with a decelerating phase. | none | Fits most tumour growth series | Wikipedia | none | ||
Log-kill hypothesis (Skipper) A dose of chemotherapy kills a constant fraction of cancer cells, not a constant number, so each cycle removes the same proportion, which is why treatment continues after the tumour has disappeared from scans. | none | Cell kill per dose is a constant logarithm: surviving fraction S = exp(-k·dose) independent of the starting number, so cure requires enough cycles to pass below one cell. | none | Explains why cycles continue after remission | source | none | ||
Mathematical models of cancer (mathematical oncology) Mathematical oncology writes down how tumours grow, evolve, respond to treatment and interact with the immune system as equations or simulations, then uses them to design doses, schedules and trials. The models themselves are records, each with what it was fitted to and what it was used to decide. | none | Describe the tumour, the treatment and the host as variables that change over time; fit the model to data; use it to predict what a different dose, schedule or combination would do. | none | Turns scattered observations into testable predictions | Wikipedia | none | ||
Body-surface-area dosing Chemotherapy doses are usually written per square metre of body surface, a convention from 1958 that scales drug clearance between species and people; it is imprecise, and for many newer drugs flat or weight-based doses have replaced it. | none | Dose = dose per square metre × body surface area, with surface area from height and weight; assumes clearance scales with surface area. | none | Universal convention for cytotoxics | Wikipedia | none | ||
Tumour volume doubling time How long a tumour takes to double in volume, measured from two scans; it separates cancers from benign nodules in lung screening, sorts aggressive from indolent disease and estimates how long a tumour has been present. | none | Doubling time = t·ln2 / ln(V2/V1) for volumes V1 and V2 measured t days apart, assuming exponential growth over the interval. | none | Simple and used in screening pathways | source | none | ||
Metastatic seeding and dormancy models From Paget's seed-and-soil idea to models that estimate when metastases were seeded from a primary and how long they lay dormant, these frameworks explain late relapse and argue for treating micrometastases early. | none | Metastatic burden follows from a seeding rate proportional to primary size and the growth law of secondaries, with dormancy as a quiescent or immune-controlled state that can be released. | none | Explains late relapse and organ tropism | Wikipedia | none | ||
Cancer Models (PDCM Finder) & HCMI Find a mouse or dish model that matches a tumour type or mutation. | none | Thousands of PDX, organoid and cell-line models across providers | Model metadata open; models by request from providers. | portal | none | none | ||
cBioPortal for Cancer Genomics cBioPortal lets you browse the mutations, copy number, and expression of tens of thousands of tumours without writing code. | none | More than 300 cancer genomics studies | Per-study; public studies are de-identified. | portal | none | none | ||
CIViC CIViC is an open, Wikipedia-style database of what cancer mutations mean for treatment. | none | Thousands of curated clinical variant interpretations | CC0. | portal | none | none | ||
COSMIC (Catalogue of Somatic Mutations in Cancer) COSMIC, the Catalogue of Somatic Mutations in Cancer, is the largest curated catalogue of mutations found in tumours, run by the Wellcome Sanger Institute. It also holds the Cancer Gene Census of around 750 genes, the reference mutational signatures and a cell line project; academic use is free, industry needs a licence. | none | Curated somatic mutations across millions of samples; Cancer Gene Census of about 750 genes | Free academic registration; commercial licence. | portal | none | none | ||
CPTAC (Clinical Proteomic Tumor Analysis Consortium) CPTAC measures the proteins, not just the genes, of thousands of tumours. | none | Proteogenomic profiles on more than 1,000 TCGA-linked tumours | Open processed data; controlled raw data through dbGaP. | portal | CHIEF | none | ||
CZ CELLxGENE / Human Cell Atlas CELLxGENE and the Human Cell Atlas hold single-cell data from healthy and diseased tissue, browsable and downloadable. | none | Tens of millions of annotated single cells | CC BY; donor consent handled by contributing studies. | portal | scGPT, Geneformer, Universal Cell Embedding | none | ||
Flatiron Health and Foundation Medicine Clinico-Genomic Database Real-world evidence at scale: what happened to patients with a given genomic profile on a given treatment. | none | More than 100,000 US patients with linked EHR outcomes and genomic profiles | De-identified under HIPAA; commercial and research licences. | portal | none | none | ||
Human Protein Atlas The Human Protein Atlas shows where in the body each protein is found, which tells you whether an ADC target is safe. | none | Protein expression across tissues, cancers and single cells | CC BY-SA 4.0. | portal | none | none | ||
NCI CDAS: NLST and PLCO screening trial data NCI CDAS holds the lung screening trial images that trained Sybil and most lung-nodule AI. | none | NLST: 53,000 participants with low-dose CT; PLCO screening data | Data access request to NCI CDAS. | portal | Sybil | none | ||
NCI Imaging Data Commons (IDC) The Imaging Data Commons is TCIA in the cloud, ready for large-scale model training. | none | Cloud copy of TCIA and other collections in DICOM | Per collection; queryable with BigQuery. | portal | CT-FM | none | ||
OncoKB Tells you, for a given mutation, whether an approved or investigational drug exists and how strong the evidence is. | none | Precision oncology knowledge base with FDA-recognised levels of evidence | Free academic licence; commercial licence. | portal | none | none | ||
Open Targets Platform Open Targets scores how strongly each gene is linked to each disease, with the evidence behind it. | none | Target-disease evidence for about 60,000 targets | CC0 and per-source licences. | portal | none | none | ||
Pathology AI benchmarks (CAMELYON, PANDA, TCGA slide tasks) Pathology AI benchmarks are the open challenge datasets on which every pathology model is scored: CAMELYON16 and 17 for lymph node metastasis detection, PANDA for prostate grading with 11,000 biopsies, and TCGA slide-level tasks used to compare foundation models. Licences are mostly CC BY-NC-SA or set per challenge. | none | CAMELYON16/17, PANDA (11,000 biopsies) and TCGA slide tasks | CC BY-NC-SA and per-challenge terms. | portal | UNI and CONCH, Midnight | none |
Weights status and licence were read from model cards and repositories on the date in each technology record and change often; check the linked page before relying on them. Benchmark results are the developers' own claims on their own test sets and are not comparable across rows. Prospective clinical validation is the exception, not the rule; see the open question on that point.