379 open-source projects the war on cancer runs on, in fourteen categories: 349 repositories read from the GitHub API on 2026-09-23 (licence as declared, stars, last push) and 30 project pages, package indexes and model cards fetched the same day. 277 were pushed to in the last two years. Filter by category, licence family, language, openness, last activity, technology, data source, cancer or maintainer; every chip is a filter and every name opens the project.
Most of what happens between a tumour sample and a treatment decision now runs on code anyone can read. Reads are trimmed with fastp, aligned and called with GATK, Strelka or hmftools inside nf-core/sarek; variants are annotated by Ensembl VEP and interpreted against CIViC and OncoKB; signatures come from SigProfiler, clones from PyClone and PhyloWGS, and the whole cohort is explored in cBioPortal. Imaging is read in OHIF and 3D Slicer and segmented by nnU-Net and TotalSegmentator; radiotherapy plans are researched in matRad and OpenTPS and checked against Monte Carlo from Geant4, TOPAS and GATE; slides are analysed in QuPath and increasingly by foundation models whose weights are, sometimes, released.
The pipeline engine hub walks that path step by step and says which of these projects sits at each step. This page is the inventory: what exists, who maintains it, under which licence, and how open it really is. Read the licence column before reuse: a quarter of the repositories declare custom terms or none at all, and several of the best-known models release weights only for non-commercial use.
Contribute. Missing a project, or a fact is stale? The list is one file in the repository: add an entry with the repository and what it relates to, run npx tsx scripts/fetch-open-source.ts, and the licence, stars and dates are read from the source, never typed in. Or suggest it and a maintainer will. Add the project to the Open Medical Registry as well, so the wider catalogue of open medicine has it.
| Technologies | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
AlphaFold 2 Google DeepMind (and Google Research) · since 2021 DeepMind's protein structure prediction system, code and weights released; the structures fill the AlphaFold Database. | 14.9k | Structural biology infrastructure | Google DeepMind | repository, paper | |||||
nnU-Net German Cancer Research Center (DKFZ) · since 2019 The self-configuring segmentation method from DKFZ that wins most medical segmentation challenges out of the box, including brain, kidney and liver tumour tasks. | 8.9k | AI auto-contouring and adaptive planning | German Cancer Research Center | repository | |||||
MONAI NVIDIA · since 2019 The PyTorch framework for deep learning in medical imaging, co-founded by NVIDIA and King's College London, used for tumour segmentation and detection research and products. | 8.7k | AI auto-contouring and adaptive planning, AI in radiology | NVIDIA | repository, project-monai.github.io | |||||
AlphaFold 3 Google DeepMind (and Google Research) · since 2024 DeepMind's model of proteins with ligands, nucleic acids and modifications; code is released for non-commercial use and weights by request. | 8.6k | AlphaFold 3, AI-driven drug & target discovery | Google DeepMind | repository, paper | |||||
DeepChem since 2015 A Python library that democratises deep learning for drug discovery, materials and biology. | 7k | AI-driven drug & target discovery | none | repository, deepchem.io | |||||
MedSAM University of Toronto, Bo Wang lab · since 2023 Segment Anything adapted to medical images across modalities, with released weights. | 4.4k | MedSAM / SAM-Med3D, AI auto-contouring and adaptive planning | repository, nature.com | ||||||
OHIF Viewer Open Health Imaging Foundation · since 2015 The Open Health Imaging Foundation's zero-footprint web DICOM viewer, the front end of many research imaging platforms and the NCI Imaging Data Commons. | 4.3k | CT, MRI, PET/CT | repository, docs.ohif.org, paper | ||||||
Evo 2 Arc Institute · since 2025 Arc Institute's genomic foundation model across all domains of life, with open weights. | 4.2k | Evo 2 | Arc Institute | repository, paper | |||||
Boltz MIT and Recursion · since 2024 An open biomolecular structure and affinity prediction model from MIT and Recursion, released under MIT with weights. | 4.2k | Boltz-1 / Boltz-2, AI-driven drug & target discovery | repository, paper | ||||||
RDKit since 2013 The open cheminformatics toolkit that nearly all open drug discovery code depends on for molecules, fingerprints and descriptors. | 3.6k | AI-driven drug & target discovery | none | repository | |||||
OpenFold Herbert Irving Comprehensive Cancer Center, Columbia University · since 2021 A trainable, open reproduction of AlphaFold 2 from the AlQuraishi lab, with training data released. | 3.4k | Structural biology infrastructure | Herbert Irving Comprehensive Cancer Center, Columbia University | repository, paper | |||||
RFdiffusion University of Washington, Institute for Protein Design · since 2023 The Baker lab's diffusion model for designing new proteins and binders, with released weights. | 3.1k | RFdiffusion / RFdiffusion2 and ProteinMPNN, De novo designed protein binders | repository, paper | ||||||
TotalSegmentator University Hospital Basel · since 2022 Segments over a hundred anatomical structures in CT and MRI in one command, widely used for organs at risk and body composition. | 3k | AI auto-contouring and adaptive planning, CT, CT body composition and sarcopenia measurement | repository, paper | ||||||
ESM3 and ESM C EvolutionaryScale · since 2024 EvolutionaryScale's protein language models; small models have open weights, larger ones are under a non-commercial licence. | 3k | ESM3 | repository, paper | ||||||
Seurat New York Genome Center and NYU, Satija lab · since 2015 The R toolkit for single-cell genomics from the Satija lab, with integration, clustering and spatial support. | 2.8k | Single-cell & spatial profiling | repository | ||||||
3D Slicer Slicer community, led from Brigham and Women's Hospital and Kitware · since 2020 The desktop platform for medical image analysis, segmentation, registration and image-guided therapy, with hundreds of extensions and a large research community. | 2.6k | CT, MRI, PET/CT | repository, slicer.org, paper | ||||||
Scanpy scverse · since 2017 The Python toolkit for single-cell gene expression analysis, the centre of the scverse ecosystem used across tumour single-cell studies. | 2.6k | Single-cell & spatial profiling | repository, scanpy.scverse.org, paper | ||||||
Chemprop MIT · since 2019 MIT's message-passing neural networks for molecular property prediction, used in antibiotic and oncology screening papers. | 2.5k | AI-driven drug & target discovery | repository, chemprop.csail.mit.edu, paper | ||||||
fastp since 2017 An all-in-one preprocessor for sequencing reads (quality control, trimming, UMI handling) used at the head of many cancer pipelines. | 2.4k | Clinical NGS bioinformatics and variant interpretation | none | repository, paper | |||||
Cellpose since 2020 A generalist deep-learning cell segmentation model used widely on histology, multiplex and cell-culture images. | 2.4k | Multiplex immunofluorescence, Single-cell & spatial profiling | none | repository, huggingface.co, paper | |||||
pydicom since 2013 The Python library for reading and writing DICOM files that most imaging research code depends on. | 2.2k | CT | none | repository, pydicom.github.io, paper | |||||
AlphaGenome Google DeepMind (and Google Research) · since 2024 DeepMind's model predicting regulatory effects of DNA variants; API access is free for non-commercial use. | 2.2k | AlphaGenome | Google DeepMind | repository, alphagenomedocs.com, paper | |||||
Chai-1 Chai Discovery · since 2024 Chai Discovery's multi-modal structure prediction model; code and weights are released, with commercial use permitted under its terms. | 2k | Chai-1 / Chai-2, AI-driven drug & target discovery | Chai Discovery | repository, chaidiscovery.com, paper | |||||
GATK (with Mutect2) Broad Institute of MIT and Harvard · since 2014 The Broad Institute's Genome Analysis Toolkit: variant discovery for germline and somatic DNA, including the Mutect2 somatic caller and copy-number tools most cancer pipelines start from. | 2k | Clinical NGS bioinformatics and variant interpretation, Whole-exome & whole-genome sequencing | Broad Institute of MIT and Harvard | repository, software.broadinstitute.org | |||||
OpenMM since 2013 A high-performance molecular dynamics toolkit used for binding free energy and protein simulation in drug discovery. | 2k | Structural biology infrastructure | none | repository | |||||
Galaxy Galaxy Project · since 2015 A web platform for accessible, reproducible data analysis with thousands of tools, including full somatic variant and cancer workflows on public servers. | 1.9k | Clinical NGS bioinformatics and variant interpretation, Genomics cloud and secure research environments | repository, galaxyproject.org | ||||||
ProteinMPNN University of Washington, Institute for Protein Design · since 2022 Fast sequence design for a given protein backbone, from the Baker lab. | 1.9k | De novo designed protein binders | repository, paper | ||||||
CLAM Mahmood lab, Brigham and Women's Hospital and Harvard Medical School · since 2020 Clustering-constrained attention multiple instance learning: the weakly supervised whole-slide classification pipeline from the Mahmood lab that many pathology AI papers build on. | 1.7k | Digital pathology & AI | repository, paper | ||||||
scvi-tools scverse · since 2017 Deep probabilistic models for single-cell omics: integration, annotation and differential expression. | 1.7k | Single-cell & spatial profiling | repository, paper | ||||||
ITK Insight Software Consortium, Kitware · since 2010 The Insight Toolkit for image segmentation and registration, the foundation under 3D Slicer, ANTs and much of medical image computing. | 1.7k | CT, MRI | repository, itk.org, paper |
Closed source, source available only under a restrictive licence, not about cancer, or a fetch that could not verify the licence. Each is on record so the gap is a decision, not an oversight.
Method. The list is hand-curated in scripts/open-source-curated.ts (one record per project, category, openness and what in OnCo it relates to). scripts/fetch-open-source.ts then reads each repository from the GitHub REST API and each non-repository project from its own page, and records the licence as declared (SPDX id, or "custom" where GitHub cannot map the terms, "none stated" where there is no licence file, "not stated" where a page named none), the primary language, stars, the year the repository was created, the date of the last push and the first non-Zenodo DOI the README names. Summaries are OnCo's plain-English glosses; the project's own description travels with the record in the API as /api/v1/open-source.json. Stars measure attention, not quality; a licence read by machine is a starting point, not legal advice.
Missing a project or a link to a technology? Suggest an edit. Builders: the recipes show how to read this table from the API.