遇见数据集

Significant genes for mean corpuscular volume (MCV) in the UK Biobank analysis using gene-<i>ε</i>-EN.

收藏
NIAID Data Ecosystem2026-03-11 收录
官方服务:

资源简介:

Here, we analyze 17,680 genes from N = 349,468 individuals of European-ancestry. This file gives the gene-ε gene-level association P-values using Elastic Net regularized effect sizes when gene boundaries are defined by (page 1) using UCSC annotations directly, and (page 2) augmenting the gene boundaries by adding SNPs within a ±50kb buffer. Significance was determined by using a Bonferroni-corrected P-value threshold (in our analyses, P = 0.05/14322 autosomal genes = 3.49×10−6 and P = 0.05/17680 autosomal genes = 2.83×10−6, respectively). The columns of tables on both pages provide: (1) chromosome position; (2) gene name; (3) gene-ε-EN gene P-value; (4) gene-specific heritability estimates; (5) whether or not an association between gene and trait is listed in the GWAS catalog (marked as “yes” or “no”); (6-7) the starting and ending position of the gene’s genomic position; (8) number of SNPs within a gene that were included in analysis; (9) the most significant SNP according to GWA summary statistics; (10) the P-value of the most significant SNP; and, on the first page, (11) the corresponding gene-level posterior enrichment probability as found by RSS for comparison. Note that an “NA” in column (11) occurs wherever the MCMC for RSS failed to converge. Highlighted rows represent enriched genes whose top SNP is not marginally significant according to a genome-wide Bonferroni-corrected threshold (P = 4.67×10−8 correcting for 1,070,306 SNPs analyzed). (XLSX)

本研究针对欧洲血统的349468名个体,分析了17680个基因。本文件提供了基因ε(gene-ε)的基因水平关联P值,该值基于弹性网络(Elastic Net)正则化效应量计算,对应两种基因边界定义方式:(第1页)直接采用UCSC注释定义基因边界;(第2页)通过在±50kb范围内添加单核苷酸多态性(SNP, Single Nucleotide Polymorphism)来扩充基因边界。 显著性检验采用邦费罗尼校正(Bonferroni-corrected)后的P值阈值:本分析中,针对14322个常染色体基因的校正阈值为P=0.05/14322≈3.49×10⁻⁶,针对17680个常染色体基因的校正阈值为P=0.05/17680≈2.83×10⁻⁶,二者分别对应上述两种基因边界定义方式。 两页表格的列依次包含以下信息:(1) 染色体位置;(2) 基因名称;(3) 基因ε-弹性网络(gene-ε-EN)基因P值;(4) 基因特异性遗传力估计值;(5) 该基因与性状的关联是否已被收录至全基因组关联研究(GWAS, Genome-Wide Association Study)目录,标注为“是”或“否”;(6-7) 该基因的基因组起始与终止位置;(8) 分析中纳入的基因内SNP数量;(9) 基于全基因组关联(GWA, Genome-Wide Association)汇总统计量得到的最显著SNP;(10) 该最显著SNP的P值;此外第1页还包含第(11)列:通过残差平方和(RSS, Residual Sum of Squares)分析得到的对应基因水平后验富集概率,用于对照比较。 需注意:当RSS分析的马尔可夫链蒙特卡洛(MCMC, Markov Chain Monte Carlo)算法未收敛时,第(11)列将显示为“NA”。高亮行代表富集基因,其最显著SNP未达到全基因组邦费罗尼校正后的显著性阈值(针对本次分析的1070306个SNP,校正阈值为P=4.67×10⁻⁸)。 (XLSX)

创建时间:
2020-06-15
二维码
社区交流群
二维码
科研交流群
商业服务