遇见数据集

Rider_Dependent_Databases_v1.1

收藏
Zenodo2026-08-14 更新2026-08-20 收录
官方服务:

资源简介:

This repository contains the essential alignment database and dependency submodules required for the operation of Rider, a computational tool designed for RNA viral protein identification and analysis. This collection serves as a comprehensive resource for validating structural similarities between Rider-predicted candidates and known reference RNA viral proteins, specifically focusing on RNA-dependent RNA polymerase (RdRp). Related Publication:The preprint describing the Rider methodology and this dataset is available on bioRxiv: Title: Expanding the RNA Virus Universe by Scalable Structure-Guided Discovery (2025-11-29) Link: https://doi.org/10.1101/2025.11.24.690314 Key Components: RSDB (Rider Broad Structure Library):Contains 217,762 predicted protein structures in .pdb format. To support the structural validation of putative RNA virus proteins, this large-scale library sources sequences from multiple public datasets (NCBI, Neri et al., Zayed et al., and Edgar et al.). The database consists of two main components: (1) standard canonical RdRp core domains (primarily from Neri, Zayed, and Edgar), and (2) NCBI data comprising both standard RdRps and long viral polyproteins. Notably, due to the 1,024-residue input length limit of ESMFold, structure prediction for long polyproteins is restricted to their N-terminal regions. Consequently, if an RdRp domain is located beyond the first 1,024 residues, the predicted structure captures the regions upstream of the RdRp (e.g., helicases, proteases, and capsid proteins). Rather than filtering these out, we retained this full diversity to enable comprehensive structural comparisons and the detection of diverse viral fragments or mobile genetic elements that may lack a canonical RdRp domain but still represent valid viral signals. RDSDB (Rider RdRp-domain-specific Database):Contains 189,694 refined, domain-specific RdRp structures in .pdb format (modeled using ESMFold v1). To address the specificity challenges posed by multi-domain polyproteins, this database was constructed by rigorously scanning and extracting only the RdRp catalytic domains using HMM profiles. Extraneous non-RdRp regions were removed, making this dataset highly optimized for detecting fragmented viral sequences. RDSDB30 (Non-redundant RdRp Database):Contains 9,735 representative RdRp domain structures in .pdb format. This highly curated dataset was generated by clustering the RDSDB structures using Foldseek at a 30% sequence identity threshold. This domain-centric, redundancy-reduced database serves as a high-density, computationally efficient target for rapid structural alignments, particularly for short or fragmented metatranscriptomic contigs. Integrated Submodules: ESM2 (35M parameters): Pre-trained weights utilized for protein sequence tokenization and embedding generation. ESMFold (v1 model): Model weights for high-accuracy protein structure prediction. Foldseek: Binary/Executable components for fast and sensitive structural alignment. Usage Note: This version is released primarily to support the peer review process of the Rider manuscript. The repository will be updated with refined datasets and documentation upon the official publication of Rider.

提供机构:
Zenodo
创建时间:
2025-12-06
二维码
社区交流群
二维码
科研交流群
商业服务