Population structure and eigenanalysis.

Population structure and eigenanalysis.
复制标题

DOI:
10.1371/journal.pgen.0020190
复制
发表时间:
2006-12
期刊:
影响因子:
4.5
通讯作者:
Reich D
Reich D
中科院分区:
生物学2区
文献类型:
--
作者:
Patterson N;Price AL;Reich D

文献摘要

参考文献

被引文献

相似文献

目前从遗传数据推断群体结构的方法没有提供正式的群体分化显著性检验。我们讨论了一种研究种群结构的方法(主成分分析),这种方法首先由Cavalli-Sforza及其同事应用于遗传数据。我们将该方法建立在坚实的统计基础上,使用现代统计学的结果来开发正式的显著性检验。我们还发现了一个关于检测遗传数据中结构的能力的一般“相变”现象,这一现象出现在我们使用的统计理论中,并对发现遗传数据中结构的能力具有重要意义:对于固定但较大的数据集大小,两个总体之间的差异(例如,如通过像FST的统计测量的)低于阈值基本上是不可检测的,但是稍微高于阈值,检测将是容易的。这意味着我们可以预测检测结构所需的数据集大小。在分析遗传数据时,人们通常希望确定样本是否来自具有结构的群体。这些样本是否可以被视为从同质群体中随机选择的,或者这些数据是否意味着该群体在遗传上不是同质的?Patterson、Price和赖希表明,一种古老的方法(主成分)与现代统计学(Tracy-Widom理论)相结合,可以快速有效地回答这个问题。该技术在最大的数据集上简单实用,并且可以应用于双等位基因的遗传标记和高度多态性的标记(例如微卫星)。该理论还允许作者估计检测结构所需的数据大小,如果他们的样本实际上来自两个具有给定但较小分化水平的群体。
Current methods for inferring population structure from genetic data do not provide formal significance tests for population differentiation. We discuss an approach to studying population structure (principal components analysis) that was first applied to genetic data by Cavalli-Sforza and colleagues. We place the method on a solid statistical footing, using results from modern statistics to develop formal significance tests. We also uncover a general “phase change” phenomenon about the ability to detect structure in genetic data, which emerges from the statistical theory we use, and has an important implication for the ability to discover structure in genetic data: for a fixed but large dataset size, divergence between two populations (as measured, for example, by a statistic like FST) below a threshold is essentially undetectable, but a little above threshold, detection will be easy. This means that we can predict the dataset size needed to detect structure. When analyzing genetic data, one often wishes to determine if the samples are from a population that has structure. Can the samples be regarded as randomly chosen from a homogeneous population, or does the data imply that the population is not genetically homogeneous? Patterson, Price, and Reich show that an old method (principal components) together with modern statistics (Tracy–Widom theory) can be combined to yield a fast and effective answer to this question. The technique is simple and practical on the largest datasets, and can be applied both to genetic markers that are biallelic and to markers that are highly polymorphic such as microsatellites. The theory also allows the authors to estimate the data size needed to detect structure if their samples are in fact from two populations that have a given, but small level of differentiation.
DOI: 10.1086/368276
发表时间: 2003-03-01
影响因子: 9.8
作者:
Allen, AS;Rathouz, PJ;Satten, GA
通讯作者: Satten, GA
DOI: 10.1016/j.jmva.2005.08.003
发表时间: 2006-07-01
影响因子: 1.6
作者:
Baik, Jinho;Silverstein, Jack W.
通讯作者: Silverstein, Jack W.
DOI: 10.1038/368455a0
发表时间: 1994-03-31
期刊: NATURE
影响因子: 64.8
作者:
BOWCOCK, AM;RUIZLINARES, A;CAVALLISFORZA, LL
通讯作者: CAVALLISFORZA, LL
DOI: 10.1111/j.1529-8817.2005.00224.x
发表时间: 2006-03-01
影响因子: 1.9
作者:
Capelli, C;Redhead, N;Goldstein, DB
通讯作者: Goldstein, DB
DOI: 10.1038/ng1653
发表时间: 2005-11-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Clayton, DG;Walker, NM;Todd, JA
通讯作者: Todd, JA