Local PCA Shows How the Effect of Population Structure Differs Along the Genome

Local PCA Shows How the Effect of Population Structure Differs Along the Genome
复制标题

DOI:
10.1534/genetics.118.301747
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.3
通讯作者:
Ralph, Peter
Ralph, Peter
中科院分区:
生物学2区
文献类型:
--
作者:
Li, Han;Ralph, Peter

文献摘要

被引文献

相似文献

群体结构导致了大型基因组数据集中个体之间平均关联度的系统模式,这通常是使用主成分分析(PCA)等降维技术来发现和可视化的。平均亲缘关系是指特定于基因座的系谱树之间的关系的平均值,它可以在中间基因组尺度上受到连锁选择和其他因素的强烈影响。我们展示了如何使用局部主成分分析来描述关联模式中的这种中等规模的异质性,并将该方法应用于来自三个物种的基因组数据,发现在每个物种中,种群结构的影响可能只在几个兆基上有很大差异。在全球人类数据集中,局部化的异质性很可能是由多态染色体倒置来解释的。在元宝紫花苜蓿的全范围数据集中,产生异质性的因素在染色体之间是共享的,与局部基因密度相关,并且可能是连锁选择引起的,如背景选择或局部适应。在一个主要是非洲果蝇的数据集中,每条染色体臂上的大规模异质性可以通过已知的染色体倒置来解释,这些倒置被认为是最近选择的,在去除携带倒置的样本后,剩余的异质性与重组率和基因密度相关,这再次表明连锁选择的作用。可视化方法为发现遗传变异的生物驱动因素提供了一种灵活的新方法,它在数据中的应用突显了连锁选择和染色体倒置对观察到的遗传变异模式的强大影响。
Population structure leads to systematic patterns in measures of mean relatedness between individuals in large genomic data sets, which are often discovered and visualized using dimension reduction techniques such as principal component analysis (PCA). Mean relatedness is an average of the relationships across locus-specific genealogical trees, which can be strongly affected on intermediate genomic scales by linked selection and other factors. We show how to use local PCA to describe this intermediate-scale heterogeneity in patterns of relatedness, and apply the method to genomic data from three species, finding in each that the effect of population structure can vary substantially across only a few megabases. In a global human data set, localized heterogeneity is likely explained by polymorphic chromosomal inversions. In a range-wide data set of Medicago truncatula, factors that produce heterogeneity are shared between chromosomes, correlate with local gene density, and may be caused by linked selection, such as background selection or local adaptation. In a data set of primarily African Drosophila melanogaster, large-scale heterogeneity across each chromosome arm is explained by known chromosomal inversions thought to be under recent selection and, after removing samples carrying inversions, remaining heterogeneity is correlated with recombination rate and gene density, again suggesting a role for linked selection. The visualization method provides a flexible new way to discover biological drivers of genetic variation, and its application to data highlights the strong effects that linked selection and chromosomal inversions can have on observed patterns of genetic variation.