# Pool every immunotherapy trial's biomarker data into one commons

Source: https://onco.cc/ideas/idea-bio2-io-biomarker-data-commons/  
OnCo record `idea-bio2-io-biomarker-data-commons` (Idea). Data CC BY-NC 4.0, attribute "Data from OnCo (onco.cc)"; commercial use needs a licence.

## TL;DR

Dozens of trials have collected immune, genomic and imaging data on the same drugs. Nobody can analyse them together, so the answer stays hidden in fragments.

## Summary

Predicting checkpoint response is a small-data problem imposed by fragmentation, not by biology: individual trials have hundreds of patients, while the aggregate is tens of thousands with multimodal data. A federated commons with harmonised data models, standardised endpoints and privacy-preserving analysis, backed by a condition of funding or of approval, would allow multimodal models to be trained and, importantly, externally validated.

## Fields

- Kind: Idea
- Last checked: 2026-09-08
- Hypothesis: A pooled multimodal dataset of more than 10,000 checkpoint-treated patients yields a validated predictor that outperforms PD-L1 and tumour mutational burden by a clinically meaningful margin in prospective use.
- Rationale: Every previous jump in biological prediction followed data aggregation rather than method novelty, from genome-wide association studies to protein structure prediction. Existing single-trial models fail external validation, which is the signature of insufficient training diversity.
- Proposed test: Start with three sponsors and two academic consortia contributing harmonised data for one tumour type, and publish an externally validated model plus the harmonisation standard itself.
- Maturity: speculative
- Actor: data

## Sources

- Bottleneck evidence (No one can predict who responds to immunotherapy): Haslam & Prasad (JAMA Network Open 2019): https://doi.org/10.1001/jamanetworkopen.2019.2535

## Connected records

- ideas: [Patient-level multimodal foundation models for treatment selection](https://onco.cc/ideas/idea-multimodal-foundation-model/)
- collections: [AACR Project GENIE](https://onco.cc/collections/genie/), [cBioPortal for Cancer Genomics](https://onco.cc/collections/cbioportal/), [ClinicalTrials.gov](https://onco.cc/collections/clinicaltrials-gov/), [OncoKB](https://onco.cc/collections/oncokb/)
- technologies: [Pathology & radiology foundation models](https://onco.cc/technologies/pathology-foundation-model/), [RNA sequencing & expression profiling](https://onco.cc/technologies/rna-seq/), [Single-cell & spatial profiling](https://onco.cc/technologies/single-cell-spatial/)
- terms: [Real-world evidence](https://onco.cc/terms/real-world-evidence/), [Tumour mutational burden (TMB)](https://onco.cc/terms/tmb/), [Tumour proportion score (TPS)](https://onco.cc/terms/tps/)
- bottlenecks: [AI that is built but not validated or deployed](https://onco.cc/bottlenecks/b-ai-validation/), [Data silos](https://onco.cc/bottlenecks/b-data-silos/), [No one can predict who responds to immunotherapy](https://onco.cc/bottlenecks/b-immunotherapy-response/)
- key papers: [Estimation of the Percentage of US Patients With Cancer Who Are Eligible for and Respond to Checkpoint Inhibitor Immunotherapy Drugs](https://onco.cc/key-papers/paper-haslam-jama-netw-open/)

---
JSON: https://onco.cc/api/v1/entities/idea-bio2-io-biomarker-data-commons.json