Local SNP-explained methylation variation in the human brain: Supplementary Data
收藏资源简介:
The following large derived data tables supplement our manuscript titled Local SNP-explained methylation variation reveals genetically anchored and exposure-associated methylation architecture in the human brain. Each file is a UTF-8 tab-separated value table with a header row; column definitions accompany each release. Supplementary Data 1—Per-region cell-type proportions. Multi-subject Single-cell (MuSiC) deconvolution estimates (sample × cell type) for caudate (n=487), dorsolateral prefrontal cortex (DLPFC) (n=500), and hippocampus (n=452) bulk RNA-seq samples, used as RNA-level covariates in the local methylation–transcriptional feature association models. Supplementary Data 2—Full pathway and functional enrichment tables. rGREAT outputs across eight annotation databases (Gene Ontology Biological Process [GO:BP], GO Molecular Function [GO:MF], KEGG, Reactome, UniProt, MSigDB C7 immune signatures, and the GO:BP + KEGG combined set) for high-SNP, low-SNP, and low-prediction variably methylated regions (VMRs) in the caudate, dorsolateral prefrontal cortex (DLPFC), and hippocampus, stratified by donor group (BA, Black American; WA, non-Hispanic white American; and ancestry-matched). Supplementary Data 3–Regulatory-context replication statistics across cohorts. Eight cohort-level tables (one per analysis type) drawn from the Black American (BA) discovery cohort and matched multi-ancestry cohort, with brain region (tissue) and donor group (population ∈ BA,WA) preserved as columns: (i) proximity_group_tests (Wilcoxon and architecture-adjusted logistic regression for within-10 kb/within-100 kb proximity to nearest gene and tested percent-spliced-in [PSI] event); (ii) proximity_h2_group_summary and proximity_h2_bin_summary (per-class and per-quintile distance summaries); and (iii) h²-group and h²-bin discovery tables for the Activity-by-Contact (ABC) enhancer-link, nearest gene 250 kb, and PSI 250 kb local association models. Supplementary Data 4—Full S-LDSC coefficient tables. Partitioned heritability coefficients, standard errors, z-scores, raw P-values, and Benjamini–Hochberg FDR-adjusted q-values for every trait × tissue × SNP-explained class or quintile-annotation cell in the Black American (BA) discovery and matched multi-ancestry (BA; WA, non-Hispanic white American) analyses. Supplementary Data 5—Cross-donor-group SNP-based heritability estimates. Per-VMR BA (Black American) and WA (non-Hispanic white American) hSNP2 point estimates and cross-validated r2 values for all matched variably methylated regions (VMRs), with assigned SNP-explained class. Supplementary Data 6—Top-1% variable CpG sites per chromosome. Per-region, per-chromosome lists of the CpG sites retained after selecting the top 1% by inter-individual methylation SD, with the corresponding SD cut-off in the file header (n = 244,367 caudate; 211,168 dorsolateral prefrontal cortex [DLPFC]; 213,845 hippocampus). Supplementary Data 7—VMR genomic coordinates and SNP-based heritability estimates. Tab-separated BED-format files (one per brain region) reporting genomic coordinates (chr, start, end, VMR identifier), estimated hSNP2, cross-validated r2, number of clumped SNPs, and assigned SNP-explained class (high-SNP, low-SNP, low-prediction) for all VMRs retained for hSNP2 analysis in the Black American (BA) discovery cohort (n = 9,564 dorsolateral prefrontal cortex [DLPFC]; 9,838 hippocampus; 11,741 caudate) and multi-ancestry replication cohort (n = 9,995 DLPFC; 9,702 hippocampus; 11,915 caudate). These files enable independent reanalysis and integration with external genomic resources. Supplementary Data 8—VMR-level methylation and sample metadata. Summary tables including per-sample variably methylated region (VMR) methylation values and metadata annotations for each brain region. Supplementary Data 9—Associations between VMR methylation levels and sample metadata. Per-region, per-cohort summaries of variably methylated region (VMR) genomic coordinates and association statistics (effect sizes (betas), standard errors, t-values, raw p-values, and FDR-adjusted p-values) for each metadata variable derived from linear regression or vectorized matrix-regression models.



