Efficient computation of significance levels for multiple associations in large studies of correlated data, including genomewide association studies

Efficient computation of significance levels for multiple associations in large studies of correlated data, including genomewide association studies
复制标题

DOI:
10.1086/423738
复制
发表时间:
2004-09-01
影响因子:
9.8
通讯作者:
Koeleman, BPC
Koeleman, BPC
中科院分区:
生物学1区
文献类型:
--
作者:
Dudbridge, F;Koeleman, BPC

文献摘要

被引文献

相似文献

大型探索性研究,包括候选基因关联测试、全基因组连锁不平衡扫描和阵列表达实验,正变得越来越常见。这类研究的一个严重问题是,由于需要控制一大批测试的假阳性率,统计能力受到了损害。因为预计会有多种真实的关联,所以已经提出了结合来自最重要测试的证据的方法,作为单独调整测试的更强大的替代方案。目前,这些方法的实际应用受到对单核苷酸多态(SNP)关联数据相关性的依赖置换测试的限制。在全基因组范围内,对于标准标记板的重复探索来说,这既非常耗时又不切实际。在这里,我们通过将分析分布与组合证据的经验分布相匹配来缓解这些问题。对于固定长度的组合证据,我们拟合极值分布,对于最重要的长度,我们拟合贝塔分布。为了拟合这些分布,需要一个排列抽样的初始阶段,但它可以比简单的排列测试更快地完成,并且只需要为每一组测试进行一次,之后拟合的参数给出了面板的可重复使用的校准。我们的方法也是一种比标准排列测试更有效的替代方法。我们证明了我们的方法的准确性,并将其效率与国际HapMap联合会发布的全基因组SNP数据的置换测试的效率进行了比较。对综合证据的分析分布的估计将使这些强大的方法在大型探索性研究中得到更广泛的应用。
Large exploratory studies, including candidate-gene-association testing, genomewide linkage-disequilibrium scans, and array-expression experiments, are becoming increasingly common. A serious problem for such studies is that statistical power is compromised by the need to control the false-positive rate for a large family of tests. Because multiple true associations are anticipated, methods have been proposed that combine evidence from the most significant tests, as a more powerful alternative to individually adjusted tests. The practical application of these methods is currently limited by a reliance on permutation testing to account for the correlated nature of single-nucleotide polymorphism (SNP)-association data. On a genomewide scale, this is both very time-consuming and impractical for repeated explorations with standard marker panels. Here, we alleviate these problems by fitting analytic distributions to the empirical distribution of combined evidence. We fit extreme-value distributions for fixed lengths of combined evidence and a beta distribution for the most significant length. An initial phase of permutation sampling is required to fit these distributions, but it can be completed more quickly than a simple permutation test and need be done only once for each panel of tests, after which the fitted parameters give a reusable calibration of the panel. Our approach is also a more efficient alternative to a standard permutation test. We demonstrate the accuracy of our approach and compare its efficiency with that of permutation tests on genomewide SNP data released by the International HapMap Consortium. The estimation of analytic distributions for combined evidence will allow these powerful methods to be applied more widely in large exploratory studies.