AllTheBacteria/BacCorpus100
收藏资源简介:
BacCorpus100是一个大规模细菌基因组语料库,用于训练和评估基因组感知的基因组语言模型。它包含经过质量控制的细菌基因组,以基因组级别表示,其中预测的蛋白质编码序列存储为蛋白质翻译,基因间区域存储为DNA序列。基因组使用基因组草图技术在100%同一性下进行去重。该数据集涵盖约700万个基因组、200亿个蛋白质编码特征、160亿个基因间区域、超过15万个物种以及1万多个环境。数据集适用于细菌基因组语言模型的大规模预训练和分析。
BacCorpus100 is a large-scale bacterial genome corpus for training and evaluating genome-aware genomic language models. It contains quality-controlled bacterial genomes represented at the genome level, with predicted protein-coding sequences stored as protein translations and intergenic regions stored as DNA sequences. Genomes were deduplicated at 100% identity using genome sketching. BacCorpus100 spans approximately 7 million genomes, 20 billion protein-coding features, 16 billion intergenic regions, more than 150,000 species, and over 10,000 environments. The dataset is intended for large-scale pretraining and analysis of bacterial genomic language models.



