PR2 Analysis of Global SARS-CoV-2 Genomic Composition Across 35.3 Million Sequences
收藏资源简介:
Comprehensive dataset of nucleotide composition metrics and classification scores for 35,376,500 SARS-CoV-2 genome sequences, processed through the PR2 pipeline. Each record includes:- Sample identifier (header) with embedded geographic, lab, and temporal metadata (e.g., "USA/CA-LACPHL-AF03923/2021")- Nucleotide counts at third codon positions (a3s, t3s, g3s, c3s) - critical for studying mutational bias, evolutionary pressure, and QC filtering- PR2 classification confidence score (pr2_score) - for filtering high-confidence sequences Summary Statistics:- Total records: 35,375,835- Unique samples: 9,355,333- Year range: 2019 to 2025- Top country: USA (14,493,148 records)- Mean PR2 Score: 0.698 - indicating overall high classification confidence For analysis of:- Large-scale evolutionary analysis of SARS-CoV-2- Geographic and temporal trend mapping- Machine learning on genomic composition features- Quality control benchmarking in viral genomics Data files available on request! Open Science License: CC-BY-4.0 - please cite if used.



