遇见数据集

PR2 Analysis of Global SARS-CoV-2 Genomic Composition Across 35.3 Million Sequences

收藏
Zenodo2025-10-15 更新2026-05-26 收录
官方服务:

资源简介:

Comprehensive dataset of nucleotide composition metrics and classification scores for 35,376,500 SARS-CoV-2 genome sequences, processed through the PR2 pipeline. Each record includes:- Sample identifier (header) with embedded geographic, lab, and temporal metadata (e.g., "USA/CA-LACPHL-AF03923/2021")- Nucleotide counts at third codon positions (a3s, t3s, g3s, c3s) - critical for studying mutational bias, evolutionary pressure, and QC filtering- PR2 classification confidence score (pr2_score) - for filtering high-confidence sequences Summary Statistics:- Total records: 35,375,835- Unique samples: 9,355,333- Year range: 2019 to 2025- Top country: USA (14,493,148 records)- Mean PR2 Score: 0.698 - indicating overall high classification confidence For analysis of:- Large-scale evolutionary analysis of SARS-CoV-2- Geographic and temporal trend mapping- Machine learning on genomic composition features- Quality control benchmarking in viral genomics Data files available on request! Open Science License: CC-BY-4.0 - please cite if used.

提供机构:
Zenodo
创建时间:
2025-09-11
二维码
社区交流群
二维码
科研交流群
商业服务