Comprehensive evaluation of imputation performance in African Americans

Comprehensive evaluation of imputation performance in African Americans
复制标题

DOI:
10.1038/jhg.2012.43
复制
发表时间:
2012-07-01
影响因子:
3.5
通讯作者:
Arking, Dan E.
Arking, Dan E.
中科院分区:
生物学3区
文献类型:
--
作者:
Chanda, Pritam;Yuhki, Naoya;Arking, Dan E.

文献摘要

被引文献

相似文献

将全基因组单核苷酸多态(SNP)阵列归因于一个更大的已知SNPs参考小组已成为全基因组关联研究的标准和重要部分。然而,关于非裔美国人在不同的归因算法、参考人群(S)和所使用的参考SNP小组方面的归罪行为,人们知之甚少。来自社区动脉粥样硬化风险研究(ARIC)的3207个非裔美国人样本的全基因组SNP数据(Affymetrix 6.0)被用来系统地评估分配质量和产量。使用HapMap III的三个参考小组(ASW、YRI和CEU)和1000基因组计划(2010年6月发布的试点1 YRI,2010年8月和2011年6月发布的EUR和AFR)小组的SNP数据的几种组合,使用Mach、Impute和Beagle的分配算法进行归属。从每条染色体上直接分型的SNPs中约有10%被掩蔽,参考小组之间通用的SNPs被用两个统计指标-一致性准确率和Cohen‘s kappa(Kappa)系数来评估分配质量。这些指标对次要等位基因频率(MAF)和特定的基因型类别(次要等位基因纯合子、杂合子和主要等位基因纯合子)的依赖性进行了彻底的调查,以确定在非裔美国人中进行补偿的最佳小组和方法。此外,利用输入数据中每个被屏蔽的SNP的平均基因型值,研究了检测与模拟表型相关的SNP的能力。我们的结果表明,与传统使用的总一致性统计量相比,分成每个基因型类别后的基因型一致性和Cohen‘s kappa系数能够更好地区分分配绩效,并且这两个统计量都随着MAF的增加而改善,而与分配方法无关。我们还发现,无论使用哪种参照板,马赫和普特的表现都一样好,而且始终好于比格尔。在各种参考小组的组合中,对于HapMap III和1000基因组计划参考小组,多民族小组比只包含单一民族样本的小组具有更好的推算准确性。最新的1000基因组计划2011年6月发布的SNP数量大大高于HapMap III,其表现与最佳组合的HapMap III参考小组和1000基因组计划的以前版本一样或更好。《人类遗传学杂志》(2012年)57411-421;doi:10.1038/jhg.2012.43;2012年5月31日在线发布
Imputation of genome-wide single-nucleotide polymorphism (SNP) arrays to a larger known reference panel of SNPs has become a standard and an essential part of genome-wide association studies. However, little is known about the behavior of imputation in African Americans with respect to the different imputation algorithms, the reference population(s) and the reference SNP panels used. Genome-wide SNP data (Affymetrix 6.0) from 3207 African American samples in the Atherosclerosis Risk in Communities Study (ARIC) was used to systematically evaluate imputation quality and yield. Imputation was performed with the imputation algorithms MACH, IMPUTE and BEAGLE using several combinations of three reference panels of HapMap III (ASW, YRI and CEU) and 1000 Genomes Project (pilot 1 YRI June 2010 release, EUR and AFR August 2010 and June 2011 releases) panels with SNP data on chromosomes 18, 20 and 22. About 10% of the directly genotyped SNPs from each chromosome were masked, and SNPs common between the reference panels were used for evaluating the imputation quality using two statistical metrics-concordance accuracy and Cohen's kappa (kappa) coefficient. The dependencies of these metrics on the minor allele frequencies (MAF) and specific genotype categories (minor allele homozygotes, heterozygotes and major allele homozygotes) were thoroughly investigated to determine the best panel and method for imputation in African Americans. In addition, the power to detect imputed SNPs associated with simulated phenotypes was studied using the mean genotype of each masked SNP in the imputed data. Our results indicate that the genotype concordances after stratification into each genotype category and Cohen's kappa coefficient are considerably better equipped to differentiate imputation performance compared with the traditionally used total concordance statistic, and both statistics improved with increasing MAF irrespective of the imputation method. We also find that both MACH and IMPUTE performed equally well and consistently better than BEAGLE irrespective of the reference panel used. Of the various combinations of reference panels, for both HapMap III and 1000 Genomes Project reference panels, the multi-ethnic panels had better imputation accuracy than those containing only single ethnic samples. The most recent 1000 Genomes Project release June 2011 had substantially higher number of imputed SNPs than HapMap III and performed as well or better than the best combined HapMap III reference panels and previous releases of the 1000 Genomes Project. Journal of Human Genetics (2012) 57, 411-421; doi:10.1038/jhg.2012.43; published online 31 May 2012