dnagpt/omnigene4-cpt-corpus
收藏资源简介:
OmniGene-4 CPT语料库是一个用于OmniGene-4模型继续预训练的生物信息学数据集,总大小约96 GB。它包含多个子集:DNA序列(采样自公共基因组)、蛋白质序列(来自UniRef和LucaOne预训练池)、英文文本回放(采样自OpenWebText)、PDB氨基酸序列及其对应的Foldseek-3Di序列,以及DSSP二级结构标签。该数据集旨在支持多模态生物语言模型的训练,特别是在DNA、蛋白质和结构数据方面,训练时采用DNA:蛋白质:OpenWebText的1:1:1混合比例,并利用3Di和DSSP数据进行残基级分类任务。
The OmniGene-4 CPT corpus is a continued-pre-training dataset for the OmniGene-4 model, with a total size of approximately 96 GB. It includes multiple splits: DNA sequences sampled from public genomes, protein sequences derived from UniRef and the LucaOne pretraining pool, English text replay sampled from OpenWebText, PDB amino-acid sequences paired with Foldseek-3Di sequences, and DSSP secondary-structure labels. This dataset is designed for training multimodal bio-language models, particularly focusing on DNA, protein, and structural data, with a training mixing ratio of 1:1:1 for DNA:protein:OpenWebText, and utilizes 3Di and DSSP files for per-residue classification objectives.




