SARS-CoV-2 genomic surveillance: input snapshot and derived tables for a reproducible analysis pipeline
收藏资源简介:
The exact input snapshot and derived output tables for the covid-genomics analysis pipeline, deposited so that its figures can be reproduced. Why this deposit exists. The pipeline reads the Nextstrain open (GenBank/INSDC-derived) 100k build from a URL that carries no date and no version. It is refreshed regularly and previous contents are not retrievable, so without a snapshot nobody — including the author — can reproduce any figure in the repository. Every number in that repository corresponds to the build downloaded on 11 August 2026, which is what this record contains. Contents. metadata.tsv.xz, sequences.fasta.xz — the input snapshot, byte-identical to what the pipeline downloaded. manifest.json — SHA-256, byte size and fetch time for each input file, as recorded by the pipeline at download time. Verifying a file against this record is how you confirm you are holding the same bytes. processed-tables.tar.gz — the derived tables written to data/processed/: quality-filter cascade, cleaned metadata, variant and mutation frequency series, signature-mutation recovery and its threshold sweep, spike annotation, phylogenies, growth-advantage fits and dN/dS estimates. SHA256SUMS, README.md — checksums and a description of each file. Dataset. 62,159 quality-filtered human SARS-CoV-2 genomes collected between 2019-12-23 and 2026-07-25, spanning 80 months, 129 countries, 2,969 Pango lineages and 59 Nextstrain clades. Not included. Intermediate alignment products (2.7 GB) are omitted because they are regenerated from the inputs above by the pipeline's own steps. Licensing. The MIT licence covers the pipeline code and figures, not the genome sequences. Those were deposited in GenBank by laboratories worldwide and curated by the Nextstrain project; the build is INSDC-derived open data, redistributable with attribution. See NOTICE.md in the software repository for the full scope statement and the known-limitations list. Scope. This is a learning and portfolio project. It reproduces findings already established in the literature, and reproducing them is how the pipeline is shown to be correct. No result here should be cited as original research.



