Discriminant analysis of principal components: a new method for the analysis of genetically structured populations.

Discriminant analysis of principal components: a new method for the analysis of genetically structured populations.
复制标题

DOI:
10.1186/1471-2156-11-94
复制
发表时间:
2010-10-15
期刊:
影响因子:
2.9
通讯作者:
Balloux F
Balloux F
中科院分区:
生物学3区
文献类型:
--
作者:
Jombart T;Devillard S;Balloux F

文献摘要

参考文献

被引文献

相似文献

测序技术的巨大进步为破译自然种群在空间和时间上的组织提供了前所未有的前景。然而,生成的数据集的大小也带来了一些令人生畏的挑战。特别是,基于预定义的群体遗传学模型的贝叶斯聚类算法,如Structure或BAPS软件,可能无法应对如此史无前例的数据量。因此,需要较少使用计算机的方法。多变量分析似乎特别吸引人,因为它们专门致力于从大数据集中提取信息。遗憾的是,目前可用的多变量方法仍然缺乏研究自然种群遗传结构所需的一些基本特征。我们介绍了主成分判别分析(DAPC),这是一种用于识别和描述遗传相关个体集群的多变量方法。当缺乏群体先验信息时,DAPC使用序贯K-均值和模型选择来推断遗传聚类。我们的方法允许从遗传数据中提取丰富的信息,提供个体到群体的分配,群体间分化的可视评估,以及个体等位基因对群体结构的贡献。我们使用模拟数据对该方法的性能进行了评估,并以Structure为基准进行了分析。此外,我们通过分析全球人群中的微卫星多态性和季节性流感的血凝素基因序列变异来说明该方法。对模拟数据的分析表明,我们的方法在刻画种群细分方面总体上比结构更好。在DAPC中实施的用于识别集群和组间结构的图形表示的工具允许解开复杂的种群结构。我们的方法也比贝叶斯聚类算法快几个数量级,并且可能适用于更广泛的数据集。
The dramatic progress in sequencing technologies offers unprecedented prospects for deciphering the organization of natural populations in space and time. However, the size of the datasets generated also poses some daunting challenges. In particular, Bayesian clustering algorithms based on pre-defined population genetics models such as the STRUCTURE or BAPS software may not be able to cope with this unprecedented amount of data. Thus, there is a need for less computer-intensive approaches. Multivariate analyses seem particularly appealing as they are specifically devoted to extracting information from large datasets. Unfortunately, currently available multivariate methods still lack some essential features needed to study the genetic structure of natural populations. We introduce the Discriminant Analysis of Principal Components (DAPC), a multivariate method designed to identify and describe clusters of genetically related individuals. When group priors are lacking, DAPC uses sequential K-means and model selection to infer genetic clusters. Our approach allows extracting rich information from genetic data, providing assignment of individuals to groups, a visual assessment of between-population differentiation, and contribution of individual alleles to population structuring. We evaluate the performance of our method using simulated data, which were also analyzed using STRUCTURE as a benchmark. Additionally, we illustrate the method by analyzing microsatellite polymorphism in worldwide human populations and hemagglutinin gene sequence variation in seasonal influenza. Analysis of simulated data revealed that our approach performs generally better than STRUCTURE at characterizing population subdivision. The tools implemented in DAPC for the identification of clusters and graphical representation of between-group structures allow to unravel complex population structures. Our approach is also faster than Bayesian clustering algorithms by several orders of magnitude, and may be applicable to a wider range of datasets.
DOI: 10.1093/nar/gkq1079
发表时间: 2011-01
影响因子: 14.9
作者:
Benson DA;Karsch-Mizrachi I;Lipman DJ;Ostell J;Sayers EW
通讯作者: Sayers EW
DOI: 10.1093/bioinformatics/btq166
发表时间: 2010-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Kembel, Steven W.;Cowan, Peter D.;Webb, Campbell O.
通讯作者: Webb, Campbell O.
DOI: 10.1038/hdy.2008.34
发表时间: 2008-07-01
期刊: HEREDITY
影响因子: 3.8
作者:
Jombart, T.;Devillard, S.;Pontier, D.
通讯作者: Pontier, D.
DOI: 10.1098/rspb.2009.1473
发表时间: 2010-01-07
影响因子: 4.7
作者:
Amos, W.;Hoffman, J. I.
通讯作者: Hoffman, J. I.
DOI: 10.1093/bioinformatics/btq292
发表时间: 2010-08-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Jombart, Thibaut;Balloux, Francois;Dray, Stephane
通讯作者: Dray, Stephane