QC Audit of 2.5M SARS-CoV-2 RBD Sequences Reveals Systemic Ambiguity & Outliers
收藏资源简介:
This dataset and accompanying report document a quality control audit of 2,514,009 publicly deposited SARS-CoV-2 Receptor Binding Domain (RBD) nucleotide sequences. Key findings include: 42.4 million ambiguous bases (N) across the dataset (1.24% total composition)One sequence mislabeled as “RBD” with length 1.57 billion nt (longer than human genome)151,209 identical copies of an “NNNNN...” placeholder sequenceMedian length = 739 nt (expected), but extreme outliers suggest pipeline/coordinate errors Datasources: NCBI / GISAID Includes: Full summary report (TXT)Filtered clean FASTA (post-QC)Per-file counts (CSV)Reproducible Colab Python scriptsUploaded privately for portfolio/job application purposes. Available upon request. Not yet peer-reviewed - open to critique and collaboration. #SARSCoV2 #RBD #GenomicSurveillance #Bioinformatics #QualityControl #OpenScience #VariantTracking #DataAudit



