Detecting Genomic Signatures of Natural Selection with Principal Component Analysis: Application to the 1000 Genomes Data.

Detecting Genomic Signatures of Natural Selection with Principal Component Analysis: Application to the 1000 Genomes Data.
复制标题

DOI:
10.1093/molbev/msv334
复制
发表时间:
2016-04
影响因子:
10.7
通讯作者:
Blum MG
Blum MG
中科院分区:
生物学1区
文献类型:
--
作者:
Duforet-Frebourg N;Luu K;Laval G;Bazin E;Blum MG

文献摘要

被引文献

相似文献

为了描述自然选择的特征,已经开发了各种分析方法来检测候选基因组区域。我们建议使用主成分分析(PCA)对自然选择进行全基因组扫描。我们发现,种群间遗传分化的共同FST指数可以看作是主成分解释的方差比例。考虑到遗传变异和每个主成分之间的相关性,提供了一个概念框架来检测参与局部适应的遗传变异,而不需要对种群进行任何事先定义。为了验证基于PCA的方法,我们考虑了1000个基因组数据(阶段1),考虑了来自非洲、亚洲和欧洲的850个个体。在低覆盖率测序深度(3×)下获得的遗传变异数量约为3600万。遗传变异与各主成分的相关性为正选择提供了已知的靶标(EDAR、SLC24A5、SLC45A2、DARC),也提供了新的候选基因(APPBPP2、TP1A1、RTTN、KCNMA、MYO5C)和非编码RNA。除了识别参与生物适应的基因外,我们还识别了两条参与多基因适应的生物途径,这两条途径与先天免疫系统(β防御素)和脂肪代谢(脂肪酸omega氧化)有关。对欧洲数据的另一项分析表明,即使在没有明确定义的种群的情况下,基于PCA的基因组扫描也可以检索到局部适应的经典例子。基于PCA的统计在PCAdapt R包和PCAdapt快速开源软件中实现,检索众所周知的人类适应信号,这对未来的全基因组测序项目是令人鼓舞的,特别是在定义种群困难的情况下。
To characterize natural selection, various analytical methods for detecting candidate genomic regions have been developed. We propose to perform genome-wide scans of natural selection using principal component analysis (PCA). We show that the common FST index of genetic differentiation between populations can be viewed as the proportion of variance explained by the principal components. Considering the correlations between genetic variants and each principal component provides a conceptual framework to detect genetic variants involved in local adaptation without any prior definition of populations. To validate the PCA-based approach, we consider the 1000 Genomes data (phase 1) considering 850 individuals coming from Africa, Asia, and Europe. The number of genetic variants is of the order of 36 millions obtained with a low-coverage sequencing depth (3×). The correlations between genetic variation and each principal component provide well-known targets for positive selection (EDAR, SLC24A5, SLC45A2, DARC), and also new candidate genes (APPBPP2, TP1A1, RTTN, KCNMA, MYO5C) and noncoding RNAs. In addition to identifying genes involved in biological adaptation, we identify two biological pathways involved in polygenic adaptation that are related to the innate immune system (beta defensins) and to lipid metabolism (fatty acid omega oxidation). An additional analysis of European data shows that a genome scan based on PCA retrieves classical examples of local adaptation even when there are no well-defined populations. PCA-based statistics, implemented in the PCAdapt R package and the PCAdapt fast open-source software, retrieve well-known signals of human adaptation, which is encouraging for future whole-genome sequencing project, especially when defining populations is difficult.