Inferring whole-genome histories in large population datasets

Inferring whole-genome histories in large population datasets
复制标题

DOI:
10.1038/s41588-019-0483-y
复制
发表时间:
2019-09-01
期刊:
影响因子:
30.8
通讯作者:
McVean, Gil
McVean, Gil
中科院分区:
生物学1区
文献类型:
--
作者:
Kelleher, Jerome;Wong, Yan;McVean, Gil

文献摘要

被引文献

相似文献

推断一组DNA序列的完整谱系历史是进化生物学的核心问题,因为这段历史编码了影响物种的事件和力量的信息。然而,目前的方法是有限的,最精确的技术能够处理不超过一百个样品。由于现在正在收集由数百万个基因组组成的数据集,因此需要可扩展且有效的推理方法来充分利用这些资源。在这里,我们介绍了一种算法,不仅能够推断全基因组的历史与国家的最先进的准确性,但也处理四个数量级以上的序列。该方法还提供了数据的“进化编码”,从而能够有效地计算相关统计数据。我们将该方法应用于1000个基因组计划,Simons基因组多样性计划和英国生物银行的人类数据,表明推断的家谱是丰富的生物信号和有效的处理。
Inferring the full genealogical history of a set of DNA sequences is a core problem in evolutionary biology, because this history encodes information about the events and forces that have influenced a species. However, current methods are limited, and the most accurate techniques are able to process no more than a hundred samples. As datasets that consist of millions of genomes are now being collected, there is a need for scalable and efficient inference methods to fully utilize these resources. Here we introduce an algorithm that is able to not only infer whole-genome histories with comparable accuracy to the stateof-the-art but also process four orders of magnitude more sequences. The approach also provides an 'evolutionary encoding' of the data, enabling efficient calculation of relevant statistics. We apply the method to human data from the 1000 Genomes Project, Simons Genome Diversity Project and UK Biobank, showing that the inferred genealogies are rich in biological signal and efficient to process.