遇见数据集

Synthetic tutorial dataset for: Using summary data to detect and quantify ascertainment in biobanks

收藏
Zenodo2026-04-01 更新2026-05-26 收录
官方服务:

资源简介:

This repository provides a curated dataset for demonstrating and reproducing the summary-statistic–based framework for detecting and quantifying ascertainment bias in genetic studies. The data are derived from publicly available genetic resources (e.g. GWAS summary statistics and 1000 Genomes Project reference panels; see main manuscript for full details) and have been processed, harmonized, and structured for tutorial and reproducibility purposes. The dataset includes: Cohort-level allele frequencies (p_s) Census representative allele frequencies (p_r) Reference allele frequencies of the 26 populations from 1000 Genomes Project (1KGP) populations GWAS summary statistics (effect sizes) LD-pruned SNP sets used in downstream analyses These components enable users to reproduce the estimation of the ascertainment parameter (θ), including both: the standardized estimator (θ₁), and the regression-based estimator (θ₂), as described in the accompanying manuscript. The dataset is intended for: Tutorial and educational use Reproducibility of analysis workflows Method validation and benchmarking Data provenance:All underlying data originate from publicly available sources. This repository does not introduce new individual-level data. Instead, it provides processed and integrated datasets to facilitate reproducible application of the method. Users should refer to the main manuscript for full details of the original data sources and access conditions. Important:Any transformations, filtering, or restructuring applied to the dataset are for methodological demonstration and reproducibility only. This resource enables end-to-end implementation of the framework using only summary-level data, without requiring access to restricted individual-level genetic information.

提供机构:
Zenodo
创建时间:
2026-03-30
二维码
社区交流群
二维码
科研交流群
商业服务