Analysis of population structure: a unifying framework and novel methods based on sparse factor analysis.
Analysis of population structure: a unifying framework and novel methods based on sparse factor analysis.
复制标题
DOI:
10.1371/journal.pgen.1001117
复制
发表时间:
2010-09-16
期刊:
影响因子:
4.5
通讯作者:
Stephens M
中科院分区:
文献类型:
--
作者:
Engelhardt BE;Stephens M
We consider the statistical analysis of population structure using genetic data. We show how the two most widely used approaches to modeling population structure, admixture-based models and principal components analysis (PCA), can be viewed within a single unifying framework of matrix factorization. Specifically, they can both be interpreted as approximating an observed genotype matrix by a product of two lower-rank matrices, but with different constraints or prior distributions on these lower-rank matrices. This opens the door to a large range of possible approaches to analyzing population structure, by considering other constraints or priors. In this paper, we introduce one such novel approach, based on sparse factor analysis (SFA). We investigate the effects of the different types of constraint in several real and simulated data sets. We find that SFA produces similar results to admixture-based models when the samples are descended from a few well-differentiated ancestral populations and can recapitulate the results of PCA when the population structure is more “continuous,” as in isolation-by-distance models. Two different approaches have become widely used in the analysis of population structure: admixture-based models and principal components analysis (PCA). In admixture-based models each individual is assumed to have inherited some proportion of its ancestry from one of several distinct populations. PCA projects the individuals into a low-dimensional subspace. On the face of it, these methods seem to have little in common. Here we show how in fact both of these methods can be viewed within a single unifying framework. This viewpoint should help practitioners to better interpret and contrast the results from these methods in real data applications. It also provides a springboard to the development of novel approaches to this problem. We introduce one such novel approach, based on sparse factor analysis, which has elements in common with both admixture-based models and PCA. As we illustrate here, in some settings sparse factor analysis may provide more interpretable results than either admixture-based models or PCA.
登录
查看更多内容
影响因子:
9.2
作者:
Lao, Oscar;Lu, Timothy T.;Kayser, Manfred
通讯作者:
Kayser, Manfred
影响因子:
64.8
作者:
Lee, DD;Seung, HS
通讯作者:
Seung, HS
影响因子:
9.8
作者:
Nelson, Matthew R.;Bryc, Katarzyna;Lail, Eric H.
通讯作者:
Lail, Eric H.
影响因子:
56.9
作者:
Parker, HG;Kim, LV;Kruglyak, L
通讯作者:
Kruglyak, L
影响因子:
4.5
作者:
Howie, Bryan N.;Donnelly, Peter;Marchini, Jonathan
通讯作者:
Marchini, Jonathan