Uses text embeddings of gene descriptions from a general LLM to represent cells, and performs surprisingly well.
GenePT is a Stanford method that represents each gene by the text embedding of its description generated by a general large language model, and represents a cell by aggregating those gene embeddings weighted by expression. The 2023 bioRxiv preprint showed that embeddings from GPT-3.5 gene summaries rival specialised single-cell foundation models on many tasks, which raised the question of how much single-cell pretraining actually adds over what is already written about genes. It is cheap to compute and interpretable, since each dimension traces back to text, and is used as a strong baseline in benchmark studies. It does not model perturbations, so it cannot predict how a cell responds to a drug or knockout. For a newcomer: GenePT shows that a language model's knowledge of genes gets you surprisingly far in single-cell biology.
GenePT aggregates LLM gene-description embeddings weighted by expression.
Query for this technology: (TITLE:"GenePT" OR ABSTRACT:"GenePT") AND (cancer OR tumor OR tumour OR oncology OR carcinoma OR lymphoma OR leukemia OR leukaemia OR myeloma OR sarcoma OR melanoma OR glioma). Results are unfiltered search hits about GenePT, not a curated reading list.
Shares Stanford Health Care / Stanford Cancer Institute and the tags foundation-model, virtual-cell.
Shares the tags foundation-model, virtual-cell.
Shares the tags foundation-model, virtual-cell.
Shares the tags foundation-model, virtual-cell.
Shares the tags foundation-model, virtual-cell.
Shares the tags foundation-model, virtual-cell.
Shares the tags foundation-model, virtual-cell.
Shares the tags foundation-model, virtual-cell.