{"entity":{"id":"genept","kind":"technology","name":"GenePT","aka":[],"tldr":"Uses text embeddings of gene descriptions from a general LLM to represent cells, and performs surprisingly well.","summary":"GenePT is a Stanford method that represents each gene by the text embedding of its description generated by a general large language model, and represents a cell by aggregating those gene embeddings weighted by expression. The 2023 bioRxiv preprint showed that embeddings from GPT-3.5 gene summaries rival specialised single-cell foundation models on many tasks, which raised the question of how much single-cell pretraining actually adds over what is already written about genes. It is cheap to compute and interpretable, since each dimension traces back to text, and is used as a strong baseline in benchmark studies. It does not model perturbations, so it cannot predict how a cell responds to a drug or knockout. For a newcomer: GenePT shows that a language model's knowledge of genes gets you surprisingly far in single-cell biology.","status":"emerging","asOf":"2026-09-08","links":[{"label":"bioRxiv 2023","url":"https://www.biorxiv.org/content/10.1101/2023.10.16.562533v1"}],"tags":["foundation-model","virtual-cell"],"related":[],"cancers":[],"sections":["ai-computation"],"technologies":[],"targets":[],"drugs":[],"companies":[],"institutions":["stanford"],"pathways":[],"terms":[],"trials":[],"people":[],"bottlenecks":[],"keyPapers":[],"journals":[],"dependsOn":[],"notes":[],"principle":"GenePT aggregates LLM gene-description embeddings weighted by expression.","strengths":["Cheap, interpretable"],"limitations":["No perturbation modelling"],"since":2023},"route":"/technologies/genept/","neighbours":{"section":[{"id":"ai-computation","kind":"section","name":"AI & Computation","route":"/fronts/ai-computation/"}],"institution":[{"id":"stanford","kind":"institution","name":"Stanford Health Care / Stanford Cancer Institute","route":"/institutions/stanford/"}]}}