PolyploidChrBench v2.0.2: chromosome-scale ground-truth benchmark data spanning diploid-to-octoploid states, founder divergence, and population heterogeneity
收藏资源简介:
PolyploidChrBench v2.0.2 is a reproducible chromosome-scale benchmark resource designed for evaluating analytical methods under controlled polyploid complexity while retaining realistic sequence and population heterogeneity. The resource uses two independent chromosome-scale panels based on Arabidopsis thaliana Col-CEN v1.2 chromosome backbones (Chr1 and Chr3) and provides complete ground truth for ploidy, founder origin, copy composition, population assignment, sequence divergence, standing variation, haplotype structure, and private variation. The canonical design spans all integer ploidy levels from diploid to octoploid (2x–8x) and nine founder-divergence stages. S0 is a single-founder baseline, S1 is a dual-founder condition with zero fixed founder divergence, and S2–S8 form a nested dual-founder divergence series ranging from 0.025% to 2%. This design separates the transition from one to two founder sources from the subsequent increase in founder sequence divergence and therefore avoids treating these two biological dimensions as a single confounded gradient. Each ploidy-by-stage condition contains 400 chromosome-panel records distributed equally among four populations (P1–P4; 100 records per population). Across 63 conditions and two chromosome panels, the resource contains 50,400 chromosome-panel records and 252,000 chromosome-copy records. Within-population heterogeneity is introduced through four explicitly tracked layers: population-structured fixed variants, segregating standing variants, deterministic haplotype mosaics, and copy-specific private variants. The resource also provides nested sampling subsets of N40, N100, N200, and N400 per condition, allowing sample-size effects to be evaluated without generating independent population realizations. The archive contains project-generated truth tables, variant tables, copy-origin records, haplotype assignments, sample-subset definitions, configuration files, data dictionaries, technical-validation outputs, reproducibility information, and the complete workflow required to regenerate derived sequence products. Technical validation confirmed all 29 core validation criteria independently for both Chr1 and Chr3. Six representative sequence materializations and six paired-end FASTQ examples were additionally generated and validated during the canonical run. Their expected file names, sizes, read statistics, and SHA-256 checksums are retained in the archive so that regenerated outputs can be verified at the byte level. To maintain a clear licensing boundary, the upstream Col-CEN v1.2 nucleotide reference sequence and sequence-derived canonical FASTA/FASTQ files are not redistributed in this Zenodo archive. Users obtain the reference sequence independently from its upstream source and verify the exact input using the chromosome lengths and SHA-256 checksums supplied with this resource. The deterministic workflow can then regenerate the validated sequence and FASTQ examples. A fully synthetic toy reference is included for software testing and workflow verification. Licensing: Project-generated benchmark data, truth tables, validation outputs, metadata, configuration records, and project-authored documentation are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Workflow software, including Python scripts, the Snakemake workflow, shell scripts, and software tests, is licensed under the MIT License. The upstream Col-CEN reference sequence is neither redistributed nor relicensed by this archive. Detailed licence scope and third-party data information are provided in LICENSE_DATA.txt, LICENSE_CODE.txt, LICENSE_SCOPE.tsv, and the accompanying third-party data notice. PolyploidChrBench is intended as a controlled ground-truth benchmark rather than a complete simulation of natural polyploid demography or meiosis. The current release does not explicitly model selection, migration, structural variation, homeologous exchange, or full multivalent meiotic behavior. The founder-divergence stages should therefore be interpreted as controlled sequence-divergence conditions rather than literal biological categories along an autopolyploid–allopolyploid continuum.



