Source data of the MultiSTAAR manuscript "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies".
收藏资源简介:
This dataset serves as the source data for Figures 2-3 and Extended Data Figures 1-2 of the MultiSTAAR manuscript titled "A statistical framework for multi-trait rare variant analysis in large-scale whole-genome sequencing studies". MultiSTAAR is a statistical framework and computationally-scalable analytical pipeline for functionally-informed multi-trait rare variant analysis in large-scale WGS studies.Figure 2. Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and ncRNA multi-trait analysis of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerids (TG) using TOPMed data (n = 61,838).Figure 3. TOPMed genetic region (2-kb sliding window) unconditional multi-trait analysis results of low-density lipoprotein cholesterol (LDL-C), high-density lipoprotein cholesterol (HDL-C) and triglycerides (TG) using TOPMed data (n = 61,838).Extended Data Figure 1. Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of fasting glucose (FG) and fasting insulin (FI) using TOPMed data (n = 21,731).Extended Data Figure 2. Manhattan plots and Q-Q plots for unconditional gene-centric coding, noncoding and genetic region (2-kb sliding window) multi-trait analysis of C-reactive protein (CRP), interleukin 6 (IL-6), lipoprotein-associated phospholipase A2 (Lp-PLA2) activity, and lipoprotein-associated phospholipase A2 (Lp-PLA2) mass using TOPMed data (n = 9,380). In addition, LDL_individual_analysis_genome.Rdata, HDL_individual_analysis_genome.Rdata, TC_individual_analysis_genome.Rdata, TG_individual_analysis_genome.Rdata are single-variant analysis summary statistics of LDL-C, HDL-C, total cholesterol (TC) and TG for all variants with minor allele count (MAC) >= 20 using TOPMed data, respectively. Specifically, the summary statistics include the following columns: CHR: Chromosome POS: Position REF: Reference allele ALT: Alternative allele ALT_AF: Alternative allele frequency MAF: Minor allele frequency N: Sample size pvalue: Score test P value pvalue_log10: Score test P value (-log10 scale) Score: Score statistic Score_se: Standard error of the score statistic Est: Estimated effect size of the minor allele Est_se: Standard error of the estimated effect size



