Analysis of population structure: a unifying framework and novel methods based on sparse factor analysis.

Analysis of population structure: a unifying framework and novel methods based on sparse factor analysis.
复制标题

DOI:
10.1371/journal.pgen.1001117
复制
发表时间:
2010-09-16
期刊:
影响因子:
4.5
通讯作者:
Stephens M
Stephens M
中科院分区:
生物学2区
文献类型:
--
作者:
Engelhardt BE;Stephens M

文献摘要

参考文献

被引文献

相似文献

我们考虑使用遗传数据的群体结构的统计分析。我们展示了如何在一个统一的矩阵分解框架内看待两种最广泛使用的建模群体结构的方法,混合模型和主成分分析(PCA)。具体地说,它们都可以被解释为通过两个较低秩矩阵的乘积来近似观察到的基因型矩阵,但是在这些较低秩矩阵上具有不同的约束或先验分布。这为通过考虑其他约束或先验来分析人口结构的大量可能方法打开了大门。在本文中,我们介绍了这样一种新的方法,基于稀疏因子分析(SFA)。我们调查的影响,不同类型的约束在几个真实的和模拟数据集。我们发现,SFA产生类似的结果混合为基础的模型时,样本是从几个分化良好的祖先种群的后裔,可以概括PCA的结果时,人口结构是更“连续”的,隔离距离模型。两种不同的方法已成为广泛使用的人口结构分析:混合为基础的模型和主成分分析(PCA)。在基于混合的模型中,每个个体都被假设从几个不同的种群中继承了一定比例的祖先。PCA将个体投影到低维子空间中。从表面上看,这些方法似乎没有什么共同之处。在这里,我们展示了如何在事实上这两种方法可以被视为在一个统一的框架。这个观点应该有助于从业者更好地解释和对比这些方法在真实的数据应用中的结果。它还为开发解决这一问题的新方法提供了一个跳板。我们介绍了这样一种新的方法,基于稀疏因子分析,它与基于混合的模型和PCA有共同的元素。正如我们在这里所示,在某些情况下,稀疏因子分析可以提供比基于混合的模型或PCA更可解释的结果。
We consider the statistical analysis of population structure using genetic data. We show how the two most widely used approaches to modeling population structure, admixture-based models and principal components analysis (PCA), can be viewed within a single unifying framework of matrix factorization. Specifically, they can both be interpreted as approximating an observed genotype matrix by a product of two lower-rank matrices, but with different constraints or prior distributions on these lower-rank matrices. This opens the door to a large range of possible approaches to analyzing population structure, by considering other constraints or priors. In this paper, we introduce one such novel approach, based on sparse factor analysis (SFA). We investigate the effects of the different types of constraint in several real and simulated data sets. We find that SFA produces similar results to admixture-based models when the samples are descended from a few well-differentiated ancestral populations and can recapitulate the results of PCA when the population structure is more “continuous,” as in isolation-by-distance models. Two different approaches have become widely used in the analysis of population structure: admixture-based models and principal components analysis (PCA). In admixture-based models each individual is assumed to have inherited some proportion of its ancestry from one of several distinct populations. PCA projects the individuals into a low-dimensional subspace. On the face of it, these methods seem to have little in common. Here we show how in fact both of these methods can be viewed within a single unifying framework. This viewpoint should help practitioners to better interpret and contrast the results from these methods in real data applications. It also provides a springboard to the development of novel approaches to this problem. We introduce one such novel approach, based on sparse factor analysis, which has elements in common with both admixture-based models and PCA. As we illustrate here, in some settings sparse factor analysis may provide more interpretable results than either admixture-based models or PCA.
DOI: 10.1016/j.cub.2008.07.049
发表时间: 2008-08-26
期刊: CURRENT BIOLOGY
影响因子: 9.2
作者:
Lao, Oscar;Lu, Timothy T.;Kayser, Manfred
通讯作者: Kayser, Manfred
DOI: 10.1038/44565
发表时间: 1999-10-21
期刊: NATURE
影响因子: 64.8
作者:
Lee, DD;Seung, HS
通讯作者: Seung, HS
DOI: 10.1016/j.ajhg.2008.08.005
发表时间: 2008-09-12
影响因子: 9.8
作者:
Nelson, Matthew R.;Bryc, Katarzyna;Lail, Eric H.
通讯作者: Lail, Eric H.
DOI: 10.1126/science.1097406
发表时间: 2004-05-21
期刊: SCIENCE
影响因子: 56.9
作者:
Parker, HG;Kim, LV;Kruglyak, L
通讯作者: Kruglyak, L
用于下一代全基因组关联研究的灵活而准确的基因型插补方法。
DOI: 10.1371/journal.pgen.1000529
发表时间: 2009-06
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Howie, Bryan N.;Donnelly, Peter;Marchini, Jonathan
通讯作者: Marchini, Jonathan