DNA language models trained on thousands of genomes for variant and regulatory prediction.
Nucleotide Transformer from InstaDeep is a family of DNA language models trained with masked language modelling over 6-mer tokens, learning sequence patterns without labels. The Nature Methods 2024 paper describes models up to 2.5B parameters trained on 3,200 human genomes and 850 species, and fine-tunes them for regulatory element and variant effect prediction. The weights are open, which has made the models a common baseline in genomic language modelling. Their context window is short compared with newer long-context models, so they cannot see distant regulatory interactions, and their usefulness for interpreting non-coding cancer variants depends on downstream fine-tuning. For a newcomer: it is an early, freely available model that learned patterns in DNA from thousands of genomes.
Nucleotide Transformer uses masked language modelling over 6-mer tokens.
Query for this technology: (TITLE:"Nucleotide Transformer" OR ABSTRACT:"Nucleotide Transformer" OR TITLE:"InstaDeep" OR ABSTRACT:"InstaDeep") AND (cancer OR tumor OR tumour OR oncology OR carcinoma OR lymphoma OR leukemia OR leukaemia OR myeloma OR sarcoma OR melanoma OR glioma). Results are unfiltered search hits about Nucleotide Transformer (InstaDeep), not a curated reading list.
Shares Autoregressive (next-token) modelling, Genomic and protein language models: Evo 2, Enformer, ESM, Variant effect prediction and the tags foundation-model, genome.
Shares the tags foundation-model, genome.
Shares the tags foundation-model, genome.
Shares the tags foundation-model, genome.
Shares Autoregressive (next-token) modelling, Genomic and protein language models: Evo 2, Enformer, ESM, Variant effect prediction.