SAIGE-QTL: Single-nucleus Expression QTL Summary Statistics using Human Brain Cohort
收藏资源简介:
https://github.com/RajLabMSSM/SingleBrain Full associations and top association summary statistics for cis-eQTLs mapped in the human brain single-nucleus RNA seq cohort(Fujita et al.[PMID: 38514782]) using SAIGE-QTL[PMID: 38798318], as part of the "SingleBrain" project. Sample size = 402 European ancestry donors. Each file has the following naming convention: {Cell type}_eqtl_{ASSOC}.tsv.gz References The following reference was used for mapping phenotypes: 1. GENCODE - GENCODE v38 comprehensive transcripts (https://www.gencodegenes.org/human/release_38.html) Cell type The following brain 7 major cell types were tested for genetic association: Ast: astrocytes End: endothelial cells Ext: excitatory neurons IN: inhibitory neurons MG: microglia OD: oligodendrocytes OPC: oligodendrocyte progenitor cell All phenotype matrices were scaled and centered and then quantile normalized. Associations Top associations (top_assoc.tsv.gz) list the SNP-feature pair with the lowest adjusted P-value (ACAT_P) for that feature. Full associations (full_assoc.tsv.gz) list all tested SNP-feature pairs. Data dictionary The columns of the two association files only differ by the presence of the ACAT_P column in the top associations. feature: the phenotype being tested CHR: chromosome POS: position (hg38) MarkerID: the genetic variant being tested Allele1: reference allele Allele2: alternate allele AC_Allele2: Alternate allele count AF_Allele2: Alternate allele frequency in the sample MissingRate: genotype missing rate BETA: Effect size estimate (per allele effect) SE: Standard error of beta Tstat: T-statistic for the effect size var: Variance of the score test statistic p.value: P-value p.value.NA: P-value computed using a normal approximation (Wald or score test), not SPA. Is.SPA: Whether the Saddlepoint Approximation (SPA) converged (TRUE/FALSE) N: Sample size used in the test p_bonf: P-value adjusted for the number of variants tested in that feature (Bonferroni) p_FDR: P-value adjusted for the number of variants tested in that feature (FDR) ACAT_P: Aggregated Cauchy Association Test (ACAT) p-value, which is a gene-level p-value combining multiple single-variant p-values for a given gene or region.
https://github.com/RajLabMSSM/SingleBrain 本数据集为"SingleBrain"项目的组成部分,包含基于人类大脑单细胞核RNA测序队列(Fujita等[PMID: 38514782])、采用SAIGE-QTL[PMID: 38798318]分析得到的顺式表达数量性状位点(cis-eQTL)关联结果,以及其顶级关联汇总统计量。 本数据集的样本量为402名欧洲血统捐赠者。 所有文件均遵循以下命名规则:{Cell type}_eqtl_{ASSOC}.tsv.gz ### 参考文献 用于表型定位的参考数据集如下: 1. GENCODE:GENCODE v38 综合转录本(https://www.gencodegenes.org/human/release_38.html) ### 检测细胞类型 本研究针对以下7种人类大脑主要细胞类型开展遗传关联分析: - Ast:星形胶质细胞(astrocytes) - End:内皮细胞(endothelial cells) - Ext:兴奋性神经元(excitatory neurons) - IN:抑制性神经元(inhibitory neurons) - MG:小胶质细胞(microglia) - OD:少突胶质细胞(oligodendrocytes) - OPC:少突胶质前体细胞(oligodendrocyte progenitor cell) ### 数据预处理 所有表型矩阵均经过标准化中心化处理,随后进行分位数归一化。 ### 关联结果文件说明 1. 顶级关联文件(top_assoc.tsv.gz):列出每个特征对应的经校正P值(ACAT_P)最低的SNP-特征对。 2. 完整关联文件(full_assoc.tsv.gz):列出所有经过测试的SNP-特征对。 ### 数据字典 两种关联文件的列仅存在一处差异:顶级关联文件包含ACAT_P列。各字段说明如下: - feature:待检测的表型 - CHR:染色体编号 - POS:基因组位置(hg38参考基因组) - MarkerID:待检测的遗传变异位点标识 - Allele1:参考等位基因 - Allele2:变异等位基因 - AC_Allele2:变异等位基因计数 - AF_Allele2:研究样本中变异等位基因的频率 - MissingRate:基因型缺失率 - BETA:效应大小估计值(每等位基因的效应) - SE:效应估计值的标准误 - Tstat:效应大小的T统计量 - var:得分检验统计量的方差 - p.value:原始P值 - p.value.NA:基于正态近似(Wald检验或得分检验,而非鞍点近似SPA)计算得到的P值 - Is.SPA:鞍点近似(Saddlepoint Approximation, SPA)是否收敛,取值为TRUE/FALSE - N:关联检验中使用的样本量 - p_bonf:针对该特征检测的变异数进行邦费罗尼校正(Bonferroni)后的P值 - p_FDR:针对该特征检测的变异数进行错误发现率(False Discovery Rate, FDR)校正后的P值 - ACAT_P:聚合柯西关联检验(Aggregated Cauchy Association Test, ACAT)P值,即整合单个基因/区域内多个单变异P值的基因水平P值。



