Greater power and computational efficiency for kernel-based association testing of sets of genetic variants.

Greater power and computational efficiency for kernel-based association testing of sets of genetic variants.
复制标题

DOI:
10.1093/bioinformatics/btu504
复制
发表时间:
2014-11-15
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Listgarten J
Listgarten J
中科院分区:
其他
文献类型:
--
作者:
Lippert C;Xiang J;Horta D;Widmer C;Kadie C;Heckerman D;Listgarten J

文献摘要

参考文献

被引文献

相似文献

动机:基于集合的方差成分检验已被认为是一种通过聚合微弱的个体效应来增加关联性研究的力量的方法。然而,测试统计量的选择在很大程度上被忽视了,尽管它可能在获得最优功率方面发挥重要作用。我们比较了标准的统计检验--得分检验--和最近发展起来的似然比(LR)检验。此外,当需要对隐藏结构进行校正,或者寻求基因-基因交互作用时,SCORE和LR测试的最新算法在计算上可能是不切实际的。因此,我们开发了新的计算高效的方法。结果:在回顾了SCORE和LR测试之间的理论差异后,我们在真实数据上实证地发现,LR测试通常具有更大的威力。特别是,在17个真实数据集中的15个上,LR测试产生的关联至少与分数测试一样多--最多多23个关联--而在剩下的两个数据集中,分数测试最多比LR测试多产生一个关联。在合成数据上,我们发现LR测试产生的关联度高出12%,与我们在真实数据上的结果一致,但也观察到了一个极小信号的制度,其中得分测试产生的关联度比LR测试高出25%,与理论一致。最后,我们的计算加速现在支持(I)当背景核是满等级时的高效LR测试,以及(Ii)当背景核随着每次测试而改变时的高效分数测试,就像对于基因-基因相互作用测试一样。后者在规模为13 500的队列中产生了2000倍的加速比。可用性:http://research.microsoft.com/en-us/um/redmond/projects/MSCompBio/Fastlmm/.上提供的软件联系方式:herkerma@microsoft.com补充信息:补充数据可从BioInformation Online获得。
Motivation: Set-based variance component tests have been identified as a way to increase power in association studies by aggregating weak individual effects. However, the choice of test statistic has been largely ignored even though it may play an important role in obtaining optimal power. We compared a standard statistical test—a score test—with a recently developed likelihood ratio (LR) test. Further, when correction for hidden structure is needed, or gene–gene interactions are sought, state-of-the art algorithms for both the score and LR tests can be computationally impractical. Thus we develop new computationally efficient methods. Results: After reviewing theoretical differences in performance between the score and LR tests, we find empirically on real data that the LR test generally has more power. In particular, on 15 of 17 real datasets, the LR test yielded at least as many associations as the score test—up to 23 more associations—whereas the score test yielded at most one more association than the LR test in the two remaining datasets. On synthetic data, we find that the LR test yielded up to 12% more associations, consistent with our results on real data, but also observe a regime of extremely small signal where the score test yielded up to 25% more associations than the LR test, consistent with theory. Finally, our computational speedups now enable (i) efficient LR testing when the background kernel is full rank, and (ii) efficient score testing when the background kernel changes with each test, as for gene–gene interaction tests. The latter yielded a factor of 2000 speedup on a cohort of size 13 500. Availability: Software available at http://research.microsoft.com/en-us/um/redmond/projects/MSCompBio/Fastlmm/. Contact: heckerma@microsoft.com Supplementary information: Supplementary data are available at Bioinformatics online.
DOI: 10.1371/journal.pcbi.1000954
发表时间: 2010-10-14
影响因子: 4.3
作者:
Bhatia G;Bansal V;Harismendy O;Schork NJ;Topol EJ;Frazer K;Bafna V
通讯作者: Bafna V
DOI: 10.1093/biomet/86.4.929
发表时间: 1999-12-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
Kuonen, D
通讯作者: Kuonen, D
DOI: 10.1016/j.ajhg.2008.06.024
发表时间: 2008-09-12
影响因子: 9.8
作者:
Li, Bingshan;Leal, Suzanne M.
通讯作者: Leal, Suzanne M.
DOI: 10.1371/journal.pgen.1001322
发表时间: 2011-03
期刊: PLoS genetics
影响因子: 4.5
作者:
Neale BM;Rivas MA;Voight BF;Altshuler D;Devlin B;Orho-Melander M;Kathiresan S;Purcell SM;Roeder K;Daly MJ
通讯作者: Daly MJ
DOI: 10.2307/2532948
发表时间: 1995-06-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
LECESSIE, S;VANHOUWELINGEN, HC
通讯作者: VANHOUWELINGEN, HC