Tahoe-100M is the biggest single-cell dataset ever released, built to teach AI how cancer cells respond to drugs.
Tahoe-100M is the largest single-cell dataset ever released, built to teach AI models how cancer cells respond to drugs. It comprises 100 million single-cell transcriptomes across 1,100 drugs and 50 cancer cell lines, making it the biggest perturbation atlas for training virtual cell models, and it was generated with the Mosaic platform to limit batch effects. It is used to train and benchmark State (Arc Institute perturbation model) and other perturbation models. The dataset was produced by Vevo Therapeutics with Parse Biosciences and the Arc Institute and is released openly under CC BY. It is referenced by the virtual cell roadmap, the AI in oncology roadmap and the drug discovery roadmap.
Shares Arc Institute, Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell, Drug discovery roadmap: screening in mice → maps of dependency → designing in silico, AI in oncology roadmap: pattern readers → foundation models → agents in the workflow and the tags data, virtual-cell.
Shares Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell, Drug discovery roadmap: screening in mice → maps of dependency → designing in silico, AI in oncology roadmap: pattern readers → foundation models → agents in the workflow and the tag virtual-cell.
Shares Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell, Drug discovery roadmap: screening in mice → maps of dependency → designing in silico, AI in oncology roadmap: pattern readers → foundation models → agents in the workflow and the tag virtual-cell.
Shares AI in oncology roadmap: pattern readers → foundation models → agents in the workflow and the tag data.
Shares Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell and the tag virtual-cell.
Shares Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell and the tag virtual-cell.
Shares AI in oncology roadmap: pattern readers → foundation models → agents in the workflow and the tag data.
Shares Virtual cell roadmap: from bulk omics to a predictive model of a cancer cell and the tag virtual-cell.
Open-source projects that are the code behind this collection or publish it, from OnCo's own catalogue: licence and last activity as the repository reported them on the day of the fetch. Listing is not endorsement; check the licence before reuse and the validation before clinical use.
A 100 million cell single-cell perturbation atlas of 1,100 drugs across 50 cancer cell lines from Vevo Therapeutics, openly released.