遇见数据集

ViTrace Detects Viral Signatures in Tumor Transcriptomes Using a Hybrid Language Model

收藏
Zenodo2025-10-31 更新2026-05-26 收录
官方服务:

资源简介:

The first simulated dataset, Sim1, was constructed by viRNAtrap. The training and testing data are randomly sampled from the simulated reads from hg19 reference genome and 13 virus species from 7 genera across 4 phyla. ViTrace and other 5 baseline models, viRNAtrap, DeepVirFinder, ViraMiner, DeepViFi, and Seq2Seq are all trained and tested on Sim1. The second simulated dataset, Sim2, was constructed as an independent testing set to evaluate ViTrace's performance to identify viruses whose taxonomy had never been seen during the training. The third simulated dataset, Sim3, was created to evaluate ViTrace’s ability to differentiate viruses from other microorganisms. Positive samples were generated from 12,567 viral genomes, covering 8,102 species across 1,012 genera, obtained from the NCBI RefSeq database. The negative samples were derived from balanced sequences representing bacteria, fungi, and archaea. This mouse dataset originates from VVRD and NCBI GRC/m38, covering 7 microbial genera across 4 bacterial phyla. It comprises 8.0 million samples for training, 800,000 for validation, and 2.91 million for testing, and is designed for training and testing purposes.

提供机构:
Feng Zhou
创建时间:
2025-10-31
二维码
社区交流群
二维码
科研交流群
商业服务