Evaluation of genome-wide power of genetic association studies based on empirical data from the HapMap project

Evaluation of genome-wide power of genetic association studies based on empirical data from the HapMap project
复制标题

DOI:
10.1093/hmg/ddm205
复制
发表时间:
2007-10-15
影响因子:
3.5
通讯作者:
Ogawa, Seishi
Ogawa, Seishi
中科院分区:
生物学2区
文献类型:
--
作者:
Nannya, Yasuhito;Taura, Kenjiro;Ogawa, Seishi

文献摘要

被引文献

相似文献

随着高通量单核苷酸多态性(SNP)分型技术的最新进展,全基因组关联研究已成为一种现实的方法,以确定致病基因,负责复杂的遗传性状的常见疾病。在这种策略中,由于极端的多假设检验,增加的基因组覆盖率和偶然发现显示大统计量的SNP的机会之间的权衡变得严重。我们通过模拟大量的病例对照组,基于HapMap项目的经验数据,研究了这种权衡限制全基因组功率的程度。在我们的模拟中,通过经验性地计算具有增加数量的SNP的一系列标志物集合的x(2)统计量的最大值的分布来评估多个假设检验的统计成本,其用于在以下功效模拟中确定全基因组阈值。在实际的研究规模下,多重检测的成本在很大程度上抵消了增加基因组覆盖率的潜在益处,因为遗传效应适中和/或致病等位基因的频率较低。在大多数现实情况下,增加基因组覆盖率对功效的影响较小,而样本大小是全基因组关联检验可行性的主要决定因素。增加基因组覆盖率而不相应增加样本量,只会消耗资源而不会获得很少的权力。对于具有相对较大效应量的常见因果等位基因[基因型相对风险>= 1.7],我们可以预期使用现实样本量(类似于每组1000个)的现有大规模基因分型平台具有令人满意的功效。
With recent advances in high-throughput single nucleotide polymorphism ( SNP) typing technologies, genome-wide association studies have become a realistic approach to identify the causative genes that are responsible for common diseases of complex genetic traits. In this strategy, a trade-off between the increased genome coverage and a chance of finding SNPs incidentally showing a large statistics becomes serious due to extreme multiple-hypothesis testing. We investigated the extent to which this trade-off limits the genome-wide power with this approach by simulating a large number of case-control panels based on the empirical data from the HapMap Project. In our simulations, statistical costs of multiple hypothesis testing were evaluated by empirically calculating distributions of the maximum value of the x(2) statistics for a series of marker sets having increasing numbers of SNPs, which were used to determine a genome-wide threshold in the following power simulations. With a practical study size, the cost of multiple testing largely offsets the potential benefits from increased genome coverage given modest genetic effects and/or low frequencies of causal alleles. In most realistic scenarios, increasing genome coverage becomes less influential on the power, while sample size is the predominant determinant of the feasibility of genome-wide association tests. Increasing genome coverage without corresponding increase in sample size will only consume resources without little gain in power. For common causal alleles with relatively large effect sizes [genotype relative risk >= 1.7], we can expect satisfactory power with currently available large-scale genotyping platforms using realistic sample size (similar to 1000 per arm).