Efficient Coalescent Simulation and Genealogical Analysis for Large Sample Sizes.

Efficient Coalescent Simulation and Genealogical Analysis for Large Sample Sizes.
复制标题

DOI:
10.1371/journal.pcbi.1004842
复制
发表时间:
2016-05
影响因子:
4.3
通讯作者:
McVean G
McVean G
中科院分区:
生物学2区
文献类型:
--
作者:
Kelleher J;Etheridge AM;McVean G

文献摘要

被引文献

相似文献

遗传变异分析的一个核心挑战是在数百万个样本中提供逼真的基因组模拟。目前的聚结模拟没有很好地扩展,或使用的近似值,无法捕捉重要的远程链接属性。分析模拟结果也提出了一个巨大的挑战,因为目前的方法来存储家谱消耗大量的空间,解析速度慢,并没有利用相关树中的共享结构。我们通过引入稀疏树和合并记录作为系谱分析的关键单元来解决这些问题。使用这些工具,精确的模拟与重组的染色体大小的区域超过数十万个样本的合并是可能的,并大大快于目前的近似方法。我们还可以比现有方法更快地分析结果的数量级。我们对自然种群中遗传变异分布的理解是由基本生物学和人口学过程的数学模型驱动的。这种合并模型的一个关键优势是,它们能够有效地模拟我们在各种进化场景下可能看到的数据。然而,目前的方法并不适合模拟数十万样本的基因组规模数据集,如果我们要了解人口规模测序项目产生的数据,这是必不可少的。同样,处理大型模拟的结果也给研究人员带来了重大挑战,因为光是读取数据文件就可能需要很多天。在本文中,我们解决了这些问题,通过引入一种新的方式来表示信息的祖先过程。这种新的表示方法极大地提高了模拟速度和存储效率,因此大型模拟可以在几分钟内完成,输出文件可以在几秒钟内处理。
A central challenge in the analysis of genetic variation is to provide realistic genome simulation across millions of samples. Present day coalescent simulations do not scale well, or use approximations that fail to capture important long-range linkage properties. Analysing the results of simulations also presents a substantial challenge, as current methods to store genealogies consume a great deal of space, are slow to parse and do not take advantage of shared structure in correlated trees. We solve these problems by introducing sparse trees and coalescence records as the key units of genealogical analysis. Using these tools, exact simulation of the coalescent with recombination for chromosome-sized regions over hundreds of thousands of samples is possible, and substantially faster than present-day approximate methods. We can also analyse the results orders of magnitude more quickly than with existing methods. Our understanding of the distribution of genetic variation in natural populations has been driven by mathematical models of the underlying biological and demographic processes. A key strength of such coalescent models is that they enable efficient simulation of data we might see under a variety of evolutionary scenarios. However, current methods are not well suited to simulating genome-scale data sets on hundreds of thousands of samples, which is essential if we are to understand the data generated by population-scale sequencing projects. Similarly, processing the results of large simulations also presents researchers with a major challenge, as it can take many days just to read the data files. In this paper we solve these problems by introducing a new way to represent information about the ancestral process. This new representation leads to huge gains in simulation speed and storage efficiency so that large simulations complete in minutes and the output files can be processed in seconds.