HuggingFaceBio/carbon-pretraining-corpus
收藏资源简介:
Carbon预训练语料库是一个用于训练基因组基础模型(如Carbon)的数据集,包含1.73亿个DNA和RNA序列,总计1.1万亿个核苷酸。该数据集涵盖真核生物和原核生物物种的DNA和RNA序列,使用6-mer分词器时对应1800亿个token。它提供多个子集配置,包括真核生物基因组(eukaryote_generator)、信使RNA(mrna_evo2)、增强mRNA转录本(mrna_splice_evo2)、原核生物基因组(prokaryote_evo2)以及一个预采样的10B-token真核生物子集(eukaryote_generator_10B_subset),用于较小或更快的运行。数据集旨在通过处理DNA字母序列(A、T、G、C)来学习生命的统计模式,覆盖真核生物、原核生物和mRNA三个主要生物复杂性层次。
The Carbon Pretraining Corpus is a dataset intended for training genomic foundation models, such as Carbon, containing 173M DNA & RNA sequences with 1.1 trillion nucleotides. It spans DNA and RNA sequences from eukaryote and prokaryote species, totaling 180B tokens when using a 6-mer tokenizer. The dataset includes multiple configs: eukaryote genomes (eukaryote_generator), messenger RNA (mrna_evo2), augmented mRNA transcripts (mrna_splice_evo2), prokaryote genomes (prokaryote_evo2), and a pre-sampled 10B-token eukaryote subset (eukaryote_generator_10B_subset) for smaller or faster runs. It is designed to learn the statistical patterns of life by treating DNA letters (A, T, G, C) as tokens, covering three major layers of biological complexity: eukaryotes, prokaryotes, and mRNA.




