Multi-locus Analysis of Genomic Time Series Data from Experimental Evolution

Multi-locus Analysis of Genomic Time Series Data from Experimental Evolution
复制标题

DOI:
10.1101/006734
复制
发表时间:
2014-06
期刊:
影响因子:
4.5
通讯作者:
Jonathan Terhorst;C. Schlötterer;Yun S. Song
Jonathan Terhorst;C. Schlötterer;Yun S. Song
中科院分区:
生物学2区
文献类型:
--
作者:
Jonathan Terhorst;C. Schlötterer;Yun S. Song

文献摘要

被引文献

相似文献

进化和再测序(E&R)实验产生的基因组时间序列数据为了解驱动进化的机制提供了一个强大的窗口。然而,标准的群体遗传推断程序不考虑随时间连续采样,需要新的方法来充分利用现代实验进化数据。为了解决这个问题,我们开发了一个高斯过程近似的多轨迹Wright-Fisher过程的选择超过几十代的时间过程。高斯过程的平均值和协方差结构是通过计算离散时间Wright-Fisher模型中的相应矩来获得的,该模型是以一个链接的选定站点的存在为条件的。这使得我们的方法能够以近似但有原则的方式解释沿着基因组和跨采样时间点的连锁和选择的影响。使用模拟数据,我们证明了我们的方法正确地检测,定位和估计从几个连锁位点中选择的等位基因的适应度的能力。我们还研究了选择强度,初始单倍型多样性,人口规模,采样频率,实验持续时间,重复次数和测序覆盖深度的不同值,这种权力的变化。除了从实验进化数据中提供选择参数的定量估计外,我们的模型还可以被从业者用来设计具有必要功率的E&R实验。最后,我们将探讨我们的基于似然的方法可以用来推断其他模型参数,包括有效的人口规模和重组率,并讨论扩展到更复杂的模型。
Genomic time series data generated by evolve-and-resequence (E&R) experiments offer a powerful window into the mechanisms that drive evolution. However, standard population genetic inference procedures do not account for sampling serially over time, and new methods are needed to make full use of modern experimental evolution data. To address this problem, we develop a Gaussian process approximation to the multi-locus Wright-Fisher process with selection over a time course of tens of generations. The mean and covariance structure of the Gaussian process are obtained by computing the corresponding moments in discrete-time Wright-Fisher models conditioned on the presence of a linked selected site. This enables our method to account for the effects of linkage and selection, both along the genome and across sampled time points, in an approximate but principled manner. Using simulated data, we demonstrate the power of our method to correctly detect, locate and estimate the fitness of a selected allele from among several linked sites. We also study how this power changes for different values of selection strength, initial haplotypic diversity, population size, sampling frequency, experimental duration, number of replicates, and sequencing coverage depth. In addition to providing quantitative estimates of selection parameters from experimental evolution data, our model can be used by practitioners to design E&R experiments with requisite power. Finally, we explore how our likelihood-based approach can be used to infer other model parameters, including effective population size and recombination rate, and discuss extensions to more complex models.