Replication Data for: Algebraic equivalence and sampling properties of F-statistics in the Genomic Era
收藏资源简介:
This note evaluates variance- and heterozygosity-based genomic estimates for F-statistics using algebraic derivation and simulation. Starting from a binary allelic indicator variable, we decompose genetic variance into its various components among populations, among individuals within populations, and within individuals. Our derivation shows that the ANOVA-based estimate of FIS follows from expected mean squares and is consistent with the classical Wright-Cockerham framework. This estimate was tested using simulated SNP data with unequal subpopulation sizes and stochastic missing genotypes. The ANOVA-derived estimate closely recovered the parametric FIS value, whereas the simple counting estimate showed substantial upward bias, leading to approximately two-fold inflation. The convergence analysis indicated that both estimates only agree under ideal conditions, needing balanced sampling and complete genotyping. Thus, for plant breeding and conservation genetics in the field, variance-component correction is essential for reliable population-genomic inference from unbalanced genomic datasets and missing data. (2026-07-30)




