The Relationship between Imputation Error and Statistical Power in Genetic Association Studies in Diverse Populations

The Relationship between Imputation Error and Statistical Power in Genetic Association Studies in Diverse Populations
复制标题

DOI:
10.1016/j.ajhg.2009.09.017
复制
发表时间:
2009-11-13
影响因子:
9.8
通讯作者:
Rosenberg, Noah A.
Rosenberg, Noah A.
中科院分区:
生物学1区
文献类型:
--
作者:
Huang, Lucy;Wang, Chaolong;Rosenberg, Noah A.

文献摘要

被引文献

相似文献

基因分型方法为具有数百万单核苷酸多态的高分辨率全基因组关联(GWA)研究提供了一种基本技术。对于基于分配的GWA研究的优化设计和解释,重要的是要了解分配误差与检测分配标记的关联的能力之间的联系。在这里,使用2×3卡方检验,我们描述了在一个被推定的标记上达到与如果该标记上的基因型别是确定的情况下所获得的统计能力相等所需的样本大小膨胀和基因归因错误率之间的关系。令人惊讶的是,典型的推算错误率(类似于2%-6%)会导致所需样本量的大幅增加(类似于10%-60%),在一些基因类型特别难以推算的非洲人群中,所需的样本量增加高达30%-150%。在大多数群体中,估算误差每增加1%,维持权力所需的样本量就会增加5%-13%。这些结果意味着,在GWA样本量计算中,研究人员将需要考虑到即使是低水平的推定错误也可能造成相当大的权力损失,而开发额外的基因组资源以减少推定错误将转化为基于推算的检测复杂人类疾病变异所需的样本量的大幅减少。
Genotype-imputation methods provide an essential technique for high-resolution genome-wide association (GWA) studies with millions of single-nucleotide polymorphisms. For optimal design and interpretation of imputation-based GWA studies, it is important to understand the connection between imputation error and power to detect associations at imputed markers. Here, using a 2 x 3 chi-square test, we describe a relationship between genotype-imputation error rates and the sample-size inflation required for achieving statistical power at an imputed marker equal to that obtained if genotypes at the marker were known with certainty. Surprisingly, typical imputation error rates (similar to 2%-6%) lead to a large increase in the required sample size (similar to 10%-60%), and in some African populations whose genotypes are particularly difficult to impute, the required sample-size increase is as high as similar to 30%-150%. In most populations, each 1% increase in imputation error leads to an increase of similar to 5%-13% in the sample size required for maintaining power. These results imply that in GWA sample-size calculations investigators will need to account for a potentially considerable loss of power from even low levels of imputation error and that development of additional genomic resources that decrease imputation error will translate into substantial reduction in the sample sizes needed for imputation-based detection of the variants that underlie complex human diseases.