遇见数据集

ViTrace Detects Viral Signatures in Tumor Transcriptomes Using a Hybrid Language Model

收藏
Zenodo2025-10-31 更新2026-05-26 收录
官方服务:

资源简介:

The first simulated dataset, Sim1, was constructed by viRNAtrap. The training and testing data are randomly sampled from the simulated reads from hg19 reference genome and 13 virus species from 7 genera across 4 phyla. ViTrace and other 5 baseline models, viRNAtrap, DeepVirFinder, ViraMiner, DeepViFi, and Seq2Seq are all trained and tested on Sim1. The second simulated dataset, Sim2, was constructed as an independent testing set to evaluate ViTrace's performance to identify viruses whose taxonomy had never been seen during the training. The third simulated dataset, Sim3, was created to evaluate ViTrace’s ability to differentiate viruses from other microorganisms. Positive samples were generated from 12,567 viral genomes, covering 8,102 species across 1,012 genera, obtained from the NCBI RefSeq database. The negative samples were derived from balanced sequences representing bacteria, fungi, and archaea. This mouse dataset originates from VVRD and NCBI GRC/m38, covering 7 microbial genera across 4 bacterial phyla. It comprises 8.0 million samples for training, 800,000 for validation, and 2.91 million for testing, and is designed for training and testing purposes.

首个模拟数据集Sim1由viRNAtrap构建完成。其训练与测试数据均从hg19参考基因组以及跨4个门、7个属的13种病毒的模拟测序读段中随机采样得到。ViTrace及另外5种基线模型:viRNAtrap、DeepVirFinder、ViraMiner、DeepViFi与Seq2Seq,均在Sim1数据集上完成训练与测试。 第二个模拟数据集Sim2作为独立测试集构建,用于评估ViTrace识别训练阶段未接触过其分类学信息的病毒的性能。 第三个模拟数据集Sim3用于评估ViTrace区分病毒与其他微生物的能力。阳性样本从NCBI RefSeq数据库获取的12567条病毒基因组中生成,这些基因组覆盖了1012个属的8102个病毒物种。阴性样本则来自代表细菌、真菌与古菌的均衡序列集。 本小鼠数据集源自VVRD与NCBI GRC/m38,覆盖了4个细菌门下的7个微生物属。该数据集包含800万个训练样本、80万个验证样本与291万个测试样本,专为训练与测试场景打造。

提供机构:
Feng Zhou
创建时间:
2025-04-22
二维码
社区交流群
二维码
科研交流群
商业服务