Assessing statistical significance in multivariable genome wide association analysis.

Assessing statistical significance in multivariable genome wide association analysis.
复制标题

DOI:
10.1093/bioinformatics/btw128
复制
发表时间:
2016-07-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Bühlmann P
Bühlmann P
中科院分区:
其他
文献类型:
--
作者:
Buzdugan L;Kalisch M;Navarro A;Schunk D;Fehr E;Bühlmann P

文献摘要

被引文献

相似文献

动机:尽管全基因组关联研究(GWAS)对大量的单核苷酸多态性(SNP)进行基因分型,但数据通常一次分析一个SNP。单个SNP的低预测能力,加上校正多重测试所需的高显著性阈值,大大降低了GWAS的能力。结果如下:我们提出了一个程序,其中所有的SNPs分析在一个多广义线性模型,我们展示了它的使用非常高维的数据集。我们的方法产生用于评估单个SNP或SNP组的显著性的P值,同时控制所有其他SNP和家族错误率(FWER)。因此,我们的方法测试SNP是否携带关于表型的任何额外信息,超出了所有其他SNP的可用信息。这排除了可能由边缘方法产生的表型和SNP之间的虚假相关性,因为“虚假相关”的SNP只是碰巧与“真正因果”的SNP相关。此外,该方法提供了一种数据驱动的方法来识别和细化SNP组,这些SNP组共同包含关于表型的信息信号。我们证明了我们的方法的价值,将其应用到由威康信托病例控制联盟(WTCCC)分析的七种疾病。我们特别表明,我们的方法也能够找到在原始WTCCC研究中未发现的显著SNP,但在其他独立研究中重复。可用性和实施:我们研究的复制由开源Bioconductor包hierGWAS支持。联系人:peter. stat.math.ethz.ch补充信息:补充数据可在生物信息学在线获得。
Motivation: Although Genome Wide Association Studies (GWAS) genotype a very large number of single nucleotide polymorphisms (SNPs), the data are often analyzed one SNP at a time. The low predictive power of single SNPs, coupled with the high significance threshold needed to correct for multiple testing, greatly decreases the power of GWAS. Results: We propose a procedure in which all the SNPs are analyzed in a multiple generalized linear model, and we show its use for extremely high-dimensional datasets. Our method yields P-values for assessing significance of single SNPs or groups of SNPs while controlling for all other SNPs and the family wise error rate (FWER). Thus, our method tests whether or not a SNP carries any additional information about the phenotype beyond that available by all the other SNPs. This rules out spurious correlations between phenotypes and SNPs that can arise from marginal methods because the ‘spuriously correlated’ SNP merely happens to be correlated with the ‘truly causal’ SNP. In addition, the method offers a data driven approach to identifying and refining groups of SNPs that jointly contain informative signals about the phenotype. We demonstrate the value of our method by applying it to the seven diseases analyzed by the Wellcome Trust Case Control Consortium (WTCCC). We show, in particular, that our method is also capable of finding significant SNPs that were not identified in the original WTCCC study, but were replicated in other independent studies. Availability and implementation: Reproducibility of our research is supported by the open-source Bioconductor package hierGWAS. Contact: peter.buehlmann@stat.math.ethz.ch Supplementary information: Supplementary data are available at Bioinformatics online.