遇见数据集

Synthetic Genotype/Phenotype Test data: GPU-Accelerated Generalized Linear Mixed Models for Biobank-Scale Association Studies

收藏
Zenodo2026-04-28 更新2026-05-26 收录
官方服务:

资源简介:

Generalized linear mixed models are the statistical foundation of rigorous biobank-scale genetic association studies, enabling null model fitting that accounts for sample relatedness and population stratification across binary and quantitative traits.While their computational cost is substantial for a single genome-wide association study, phenome-wide analysis amplifies this burden further by requiring independent null model fits across thousands of phenotypes, making each individual model solve as fast as possible essential for scientific discovery at scale. We present a GPU-accelerated framework targeting the dominant computational bottlenecks across the full genome-wide association pipeline: streaming GPU kernels for packed genomic preprocessing, block conjugate gradient for stochastic trace estimation exploiting shared matrix structure, and blocked GPU association testing with hybrid CPU-GPU routing that preserves full statistical validity.Our contributions yield more than 10x end-to-end speedup on the Million Veteran Program dataset, one of the largest biobank datasets available, with portability validated across multiple GPU architectures. Here we present the synthetic genotype and phenotype data generated using HapGen2 (Zhan et al., 2011, Bioinformatics: https://mathgen.stats.ox.ac.uk/genetics_software/hapgen/hapgen2.html) and the 1000 Genomes Project Phase 3 reference dataset. The reference includes data from 5 ancestral populations: African (AFR), American/Latino (AMR), East Asian (EAS), European (EUR), South Asian (SAS). HapGen2 requires specific information from the reference population: - Legend files, which include SNP ID, position, A0 (non-coding allele), A1 (coding allele), SNP type (e.g. Biallelic SNP), and allele frequencies. - A haplotype file, where each row represents a SNP and every two columns represent an individual, with 0 0 representing A0 A0, 1 0 representing A1 A0, and 1 1 representing A1 A1. -Map files, which have the SNP position, combined rate (cM/Mb), and genetic map (cM). Step 1 for null model fitting - Dataset 1: $N=120,000$ and $M=100,000$ - Dataset 2: $N=400,000$ and $M=121,587$ - Genotype: Stored in \texttt{*.bed, *.fam, *.bim} files - Phenotype: Stored in \texttt{*.txt} files Step 2 for score testing - Dataset 1: $N=120,000$ and $M=148,000$ - Dataset 2: $N=400,000$ and $M=148,000$ - Genotype: Stored in \texttt{*.bgen, *.bgi} files - Phenotype: Stored in \texttt{*.txt} files

提供机构:
Zenodo
创建时间:
2026-04-28
二维码
社区交流群
二维码
科研交流群
商业服务