Scorpio Gene-Taxa Benchmark Dataset
收藏资源简介:
We used the Woltka pipeline to compile the complete Basic genome dataset, consisting of 4634 genomes, with each genus represented by a single genome. After downloading all coding sequences (CDS) from the NCBI database, we extracted 8 million distinct CDS, focusing on bacteria and archaea and excluding viruses and fungi due to inadequate gene information. To maintain accuracy, we excluded hypothetical proteins, uncharacterized proteins, and sequences without gene labels. We addressed issues with gene name inconsistencies in NCBI by keeping only genes with more than 1000 samples and ensuring each phylum had at least 350 sequences. This resulted in a curated dataset of 800,318 gene sequences from 497 genes across 2046 genera. We created four datasets to evaluate our model: a training set (Train_set), a test set (Test_set) with different samples but the same genus and gene as the training set, a Taxa_out_set excluding 18 phyla present in the training set but from different phyla, and a Gene_out_set excluding 60 genes from the training set but from the same phyla. We ensured each CDS had only one representation per genome, removing genes with multiple representations within the same species.
我们采用Woltka分析流水线(Woltka pipeline)构建完整基础基因组数据集,该数据集包含4634个基因组,每个属仅对应单一组装基因组。从美国国家生物技术信息中心(National Center for Biotechnology Information, NCBI)数据库下载所有编码序列(Coding Sequence, CDS)后,我们提取得到800万条独特的编码序列,研究范围限定于细菌与古菌,排除了病毒和真菌——因这两类的基因信息存在明显不足。 为保障数据准确性,我们剔除了假设蛋白、未表征蛋白以及无基因标记的序列。针对NCBI数据库中基因名称不一致的问题,我们仅保留样本量超过1000的基因,并确保每个门至少包含350条序列。最终得到经过质控整理的数据集,包含来自2046个属、497个基因的800318条基因序列。 我们构建了四类数据集用于模型评估:训练集(Train_set)、测试集(Test_set)——其样本与训练集不同,但所属属和基因与训练集一致;类群剔除集(Taxa_out_set)——排除训练集中包含的18个门,所有样本均来自其他不同门类;以及基因剔除集(Gene_out_set)——排除训练集中的60个基因,但所有样本均属于与训练集相同的门。我们确保每个基因组内的每条编码序列(CDS)仅存在唯一表征,剔除了同一物种内存在多条表征的基因。



