# GenePT

Source: https://onco.cc/technologies/genept/  
OnCo record `genept` (Technology). Data CC BY-NC 4.0, attribute "Data from OnCo (onco.cc)"; commercial use needs a licence.

## TL;DR

Uses text embeddings of gene descriptions from a general LLM to represent cells, and performs surprisingly well.

## Summary

GenePT is a Stanford method that represents each gene by the text embedding of its description generated by a general large language model, and represents a cell by aggregating those gene embeddings weighted by expression. The 2023 bioRxiv preprint showed that embeddings from GPT-3.5 gene summaries rival specialised single-cell foundation models on many tasks, which raised the question of how much single-cell pretraining actually adds over what is already written about genes. It is cheap to compute and interpretable, since each dimension traces back to text, and is used as a strong baseline in benchmark studies. It does not model perturbations, so it cannot predict how a cell responds to a drug or knockout. For a newcomer: GenePT shows that a language model's knowledge of genes gets you surprisingly far in single-cell biology.

## Fields

- Kind: Technology
- Status: emerging
- Last checked: 2026-09-08
- Tags: foundation-model; virtual-cell
- Principle: GenePT aggregates LLM gene-description embeddings weighted by expression.
- Since: 2023
- Strengths: Cheap, interpretable
- Limitations: No perturbation modelling

## Sources

- bioRxiv 2023: https://www.biorxiv.org/content/10.1101/2023.10.16.562533v1

## Connected records

- fronts: [AI & Computation](https://onco.cc/fronts/ai-computation/)
- institutions: [Stanford Health Care / Stanford Cancer Institute](https://onco.cc/institutions/stanford/)

---
JSON: https://onco.cc/api/v1/entities/genept.json