Investigating population stratification and admixture using eigenanalysis of dense genotypes

Investigating population stratification and admixture using eigenanalysis of dense genotypes
复制标题

DOI:
10.1038/hdy.2011.26
复制
发表时间:
2011-10-01
期刊:
影响因子:
3.8
通讯作者:
Shriner, D.
Shriner, D.
中科院分区:
生物学2区
文献类型:
--
作者:
Shriner, D.

文献摘要

被引文献

相似文献

遗传数据的主成分分析用于避免关联测试中由于使用顶级特征向量进行协变量调整进行群体分层而导致 I 类错误率膨胀,并估计独立于自我报告或种族身份的聚类或群体成员资格。特征分解将相关变量转换为相同数量的不相关变量。已经制定了许多停止规则来确定应保留哪些主要成分。随机矩阵理论的最新发展导致了顶特征值的形式假设检验,提供了另一种实现降维的方法。在本研究中,我将 Velicer 的最小平均部分检验与 EIGENSOFT 中实现的基于 Tracy-Widom 分布的检验进行比较,EIGENSOFT 是全基因组关联分析中使用最广泛的主成分分析实现。通过基于合并理论的变方差的计算机模拟,EIGENSOFT 系统地高估了重要主成分的数量。此外,混合个体样本的这种高估比未混合个体样本的高估更大。高估重要主成分的数量可能会通过调整不必要的协变量而导致关联测试中功效的损失,并可能导致对群体差异的错误推断。 Velicer 的最小平均部分检验在估计要保留的主成分数量时具有较小的偏差和较小的方差,均方误差通常为 0。 Velicer 的最小平均部分测试以 R 代码实现,适用于带或不带群体标签的全基因组基因型数据。遗传 (2011) 107, 413-420; doi:10.1038/hdy.2011.26; 2011 年 3 月 30 日在线发布
Principal components analysis of genetic data is used to avoid inflation in type I error rates in association testing due to population stratification by covariate adjustment using the top eigenvectors and to estimate cluster or group membership independent of self-reported or ethnic identities. Eigendecomposition transforms correlated variables into an equal number of uncorrelated variables. Numerous stopping rules have been developed to identify which principal components should be retained. Recent developments in random matrix theory have led to a formal hypothesis test of the top eigenvalue, providing another way to achieve dimension reduction. In this study, I compare Velicer's minimum average partial test to a test on the basis of Tracy-Widom distribution as implemented in EIGENSOFT, the most widely used implementation of principal components analysis in genome-wide association analysis. By computer simulation of vicariance on the basis of coalescent theory, EIGENSOFT systematically overestimates the number of significant principal components. Furthermore, this overestimation is larger for samples of admixed individuals than for samples of unadmixed individuals. Overestimating the number of significant principal components can potentially lead to a loss of power in association testing by adjusting for unnecessary covariates and may lead to incorrect inferences about group differentiation. Velicer's minimum average partial test is shown to have both smaller bias and smaller variance, often with a mean squared error of 0, in estimating the number of principal components to retain. Velicer's minimum average partial test is implemented in R code and is suitable for genome-wide genotype data with or without population labels. Heredity (2011) 107, 413-420; doi:10.1038/hdy.2011.26; published online 30 March 2011