遇见数据集

Source data of the manuscript "MetaSTAARlite: An all-in-one tool for biobank-scale whole-genome sequencing meta-analysis"

收藏
Zenodo2026-04-08 更新2026-05-26 收录
官方服务:

资源简介:

This dataset serves as the source data for Figure 2, Extended Data Figures 3-6, and Supplementary Figures 1-6 of the manuscript titled "MetaSTAARlite: An all-in-one tool for biobank-scale whole-genome sequencing meta-analysis". MetaSTAARlite provides a scalable and resource-efficient summary statistics-based pipeline for powerful, functionally-informed rare variant meta-analysis of biobank-scale sequencing data. The files included in this dataset are as follows: UKB_meta_TC_coding.zip: Gene-centric coding meta-analysis results of total cholesterol (TC) for a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055). UKB_meta_TC_noncoding.zip: Gene-centric noncoding meta-analysis results of total cholesterol (TC) for a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055). UKB_meta_TC_ncRNA.Rdata: Noncoding RNA (ncRNA) meta-analysis results of total cholesterol (TC) for a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055). UKB_pooled_TC_coding.zip: Gene-centric coding pooled analysis results of total cholesterol (TC) using individual-level data from the UK Biobank whole-genome sequencing data dataset ($n$ = 190,110). UKB_pooled_TC_noncoding.zip: Gene-centric noncoding pooled analysis results of total cholesterol (TC) using individual-level data from the UK Biobank whole-genome sequencing data dataset ($n$ = 190,110). UKB_pooled_TC_ncRNA.Rdata: Noncoding RNA (ncRNA) pooled analysis results of total cholesterol (TC) using individual-level data from the UK Biobank whole-genome sequencing data dataset ($n$ = 190,110). UKB_meta_TC_coding_INT.zip: Gene-centric coding sensitivity meta-analysis results of total cholesterol (TC) for a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055), in which the rank-based inverse normal transformation to the residuals were applied after the 1:2:3 random partition. UKB_meta_TC_noncoding_INT.zip: Gene-centric noncoding sensitivity meta-analysis results of total cholesterol (TC) for a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055), in which the rank-based inverse normal transformation to the residuals were applied after the 1:2:3 random partition. UKB_meta_TC_ncRNA_INT.Rdata: Noncoding RNA (ncRNA) sensitivity meta-analysis results of total cholesterol (TC) for a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055), in which the rank-based inverse normal transformation to the residuals were applied after the 1:2:3 random partition. UKB_AoU_meta_TC.zip: Gene-centric coding meta-analysis results of total cholesterol (TC) using the UK Biobank whole-exome sequencing data ($n$ = 446,933) and All of Us exome callset of the short read whole-genome sequencing data ($n$ = 94,532). UKB_AoU_meta_height.zip: Gene-centric coding meta-analysis results of height using the UK Biobank whole-exome sequencing data ($n$ = 467,038) and the All of Us exome callset of short read whole-genome sequencing data ($n$ = 222,316). UKB_AoU_meta_eGFR.zip: Gene-centric coding meta-analysis results of estimated glomerular filtration rate (eGFR) using the UK Biobank whole-exome sequencing data ($n$ = 446,314) and the All of Us exome callset of short read whole-genome sequencing data ($n$ = 22,658). UKB_AoU_meta_calcium.zip: Gene-centric coding meta-analysis results of calcium using the UK Biobank whole-exome sequencing data ($n$ = 409,114) and the All of Us exome callset of short read whole-genome sequencing data ($n$ = 129,547). UKB_AoU_meta_binary_LDL.zip: Gene-centric coding meta-analysis results of elevated low-density lipoprotein cholesterol (adjusted LDL-C > 130 mg/dL) using the UK Biobank whole-exome sequencing data ($n$ = 435,410) and All of Us exome callset of the short read whole-genome sequencing data ($n$ = 89,147). The files used as source data to each figure are as follows: Figure 2. Miami plot, quantile-quantile (Q-Q) plot, and scatterplot comparing the results obtained from gene-centric noncoding meta-analysis of total cholesterol (TC) for a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055) and those obtained from a gene-centric noncoding pooled analysis of TC using individual-level data from the same dataset ($n$ = 190,110). Source: UKB_meta_TC_noncoding.zip, UKB_pooled_TC_noncoding.zip Extended Data Figure 3. Miami plot, quantile-quantile (Q-Q) plot, and scatterplot comparing the results obtained from gene-centric coding meta-analysis of total cholesterol (TC) for a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055) and those obtained from a gene-centric coding pooled analysis of TC using individual-level data from the same dataset ($n$ = 190,110). Source: UKB_meta_TC_coding.zip, UKB_pooled_TC_coding.zip Extended Data Figure 4. Scatterplots comparing results for a 1:2:3 random partition ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055) of the UK Biobank whole-genome sequencing data that were generated by MetaSTAAR-O to results that were generated by other rare variant meta-analysis methods. Source: UKB_meta_TC_noncoding.zip, UKB_meta_TC_coding.zip Extended Data Figure 5. Manhattan plot and quantile-quantile (Q-Q) plot for meta-analysis of total cholesterol using the UK Biobank whole-exome sequencing data ($n$ = 446,933) and All of Us exome callset of the short read whole-genome sequencing data ($n$ = 94,532). Source: UKB_AoU_meta_TC.zip Extended Data Figure 6. Manhattan plot and quantile-quantile (Q-Q) plot for meta-analysis of elevated low-density lipoprotein cholesterol (adjusted LDL-C > 130 mg/dL) using the UK Biobank whole-exome sequencing data (n = 435,410) and All of Us exome callset of the short read whole-genome sequencing data (n = 89,147). Source: UKB_AoU_meta_binary_LDL.zip Supplementary Figure 1. Quantile-quantile (Q-Q) plots for gene-centric coding meta-analysis of total cholesterol (TC) and gene-centric coding pooled analysis of TC, using a 1:2:3 random partition of the UK Biobank whole-genome sequencing data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055; $n$ = 190,110). Source: UKB_meta_TC_coding.zip, UKB_pooled_TC_coding.zip Supplementary Figure 2. Quantile-quantile (Q-Q) plots for gene-centric noncoding meta-analysis of total cholesterol (TC) and gene-centric noncoding pooled analysis of TC, using a 1:2:3 random partition of the UK Biobank WGS data ($n_1$ = 31,685; $n_2$ = 63,370; $n_3$ = 95,055; $n$ = 190,110). Source: UKB_meta_TC_noncoding.zip, UKB_pooled_TC_noncoding.zip Supplementary Figure 3. Manhattan plot and quantile-quantile (Q-Q) plot for meta-analysis of height using the UK Biobank whole-exome sequencing data ($n$ = 467,038) and the All of Us exome callset of short read whole-genome sequencing data ($n$ = 222,316). Source: UKB_AoU_meta_height.zip Supplementary Figure 4. Manhattan plot and quantile-quantile (Q-Q) plot for meta-analysis of estimated glomerular filtration rate (eGFR) using the UK Biobank whole-exome sequencing data ($n$ = 446,314) and the All of Us exome callset of short read whole-genome sequencing data ($n$ = 22,658). Source: UKB_AoU_meta_eGFR.zip Supplementary Figure 5. Manhattan plot and quantile-quantile (Q-Q) plot for meta-analysis of calcium using the UK Biobank whole-exome sequencing data ($n$ = 409,114) and the All of Us exome callset of short read whole-genome sequencing data ($n$ = 129,547). Source: UKB_AoU_meta_calcium.zip Supplementary Figure 6. Scatterplots comparing gene-centric unconditional MetaSTAAR-O P values obtained from a sensitivity meta-analysis (in which the rank-based inverse normal transformation to the residuals were applied after the 1:2:3 random partition) to STAAR-O P values obtained from the joint analysis of pooled individual-level data. Source: UKB_meta_TC_coding_INT.zip, UKB_pooled_TC_coding.zip, UKB_meta_TC_noncoding_INT.zip, UKB_pooled_TC_noncoding.zip

提供机构:
Zenodo
创建时间:
2026-04-08
二维码
社区交流群
二维码
科研交流群
商业服务