Calibrating the Performance of SNP Arrays for Whole-Genome Association Studies
收藏资源简介:
To facilitate whole-genome association studies (WGAS), several high-density SNP genotyping arrays have been developed. Genetic coverage and statistical power are the primary benchmark metrics in evaluating the performance of SNP arrays. Ideally, such evaluations would be done on a SNP set and a cohort of individuals that are both independently sampled from the original SNPs and individuals used in developing the arrays. Without utilization of an independent test set, previous estimates of genetic coverage and statistical power may be subject to an overfitting bias. Additionally, the SNP arrays' statistical power in WGAS has not been systematically assessed on real traits. One robust setting for doing so is to evaluate statistical power on thousands of traits measured from a single set of individuals. In this study, 359 newly sampled Americans of European descent were genotyped using both Affymetrix 500K (Affx500K) and Illumina 650Y (Ilmn650K) SNP arrays. From these data, we were able to obtain estimates of genetic coverage, which are robust to overfitting, by constructing an independent test set from among these genotypes and individuals. Furthermore, we collected liver tissue RNA from the participants and profiled these samples on a comprehensive gene expression microarray. The RNA levels were used as a large-scale set of quantitative traits to calibrate the relative statistical power of the commercial arrays. Our genetic coverage estimates are lower than previous reports, providing evidence that previous estimates may be inflated due to overfitting. The Ilmn650K platform showed reasonable power (50% or greater) to detect SNPs associated with quantitative traits when the signal-to-noise ratio (SNR) is greater than or equal to 0.5 and the causal SNP's minor allele frequency (MAF) is greater than or equal to 20% (N = 359). In testing each of the more than 40,000 gene expression traits for association to each of the SNPs on the Ilmn650K and Affx500K arrays, we found that the Ilmn650K yielded 15% times more discoveries than the Affx500K at the same false discovery rate (FDR) level.
为推动全基因组关联研究(Whole-Genome Association Studies, WGAS)的发展,目前已开发出多款高密度单核苷酸多态性(Single Nucleotide Polymorphism, SNP)基因分型芯片。评估SNP芯片性能的核心基准指标为遗传覆盖度与统计效力。理想状态下,此类性能评估应基于独立于芯片开发所用SNP集合与研究对象队列的独立样本集开展。若未采用独立测试集,此前针对遗传覆盖度与统计效力的估算结果可能存在过拟合偏差。此外,现有研究尚未针对真实性状系统评估SNP芯片在全基因组关联研究中的统计效力。一种可靠的评估方案为:基于单一样本队列中获取的数千个性状开展统计效力评估。 本研究采用Affymetrix 500K(简称Affx500K)与Illumina 650Y(简称Ilmn650K)两款SNP基因分型芯片,对359名新招募的欧洲裔美国人进行基因分型。基于这批实验数据,我们通过从该队列的基因型数据中划分独立测试集的方式,获得了不受过拟合影响的遗传覆盖度估算结果。此外,我们还收集了参与者的肝脏组织RNA样本,并利用全覆盖基因表达微阵列对样本进行基因表达谱分析。我们将RNA表达水平作为大规模定量性状集,用于校准两款商用芯片的相对统计效力。 我们得到的遗传覆盖度估算值低于此前的研究报告,这表明此前的估算结果可能因过拟合问题被高估。Ilmn650K平台展现出合理的检测效力:当信噪比(Signal-to-Noise Ratio, SNR)≥0.5且致病变异的次要等位基因频率(Minor Allele Frequency, MAF)≥20%时,在样本量N=359的情况下,该芯片可检测到与定量性状相关联的SNP的效力可达50%及以上。 我们针对Ilmn650K与Affx500K芯片上的全部SNP,分别对超过40000个基因表达性状开展关联分析,结果显示,在相同的错误发现率(False Discovery Rate, FDR)阈值下,Ilmn650K的阳性发现数量比Affx500K高出15%。



