A coalescent-based method for population tree inference with haplotypes.

A coalescent-based method for population tree inference with haplotypes.
复制标题

一种基于合并的单倍型群体树推断方法。

DOI:
10.1093/bioinformatics/btu710
复制
发表时间:
2015
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Wu,Yufeng
Wu,Yufeng
中科院分区:
--
文献类型:
--
作者:
Wu,Yufeng

文献摘要

相似文献

动机:种群树表示过去的种群分歧历史。种群树的推断对于研究种群进化是有用的。随着大规模群体遗传项目(如1000个基因组项目)中数据量的增加,祖先群体推断(包括群体树推断)面临新的计算挑战。用于群体树推断的现有方法主要被设计用于非连锁遗传变体(例如单核苷酸多态性或SNP)。有一个潜在的损失的信息不考虑haplotypes.Results:在这篇文章中,我们提出了一个新的人口树推理方法(称为STELLSH)的基础上合并的可能性。可能性是在非重组区域内的多个SNP上的单倍型,而不是未连锁的变体。与许多现有的祖先推理方法不同,STELLSH在计算可能性时不使用Monte Carlo方法。为了提高计算效率,似然模型是近似的,但仍然保留了大量关于种群发散历史的信息。STELLSH可以根据近似的似然性找到最大似然种群树。我们通过模拟数据和1000个基因组计划的数据表明,STELLSH给出了相当准确的推理结果。STELLSH对当前感兴趣的数据相当有效,并且可以扩展到处理全基因组数据。可用性和实施:人口树推断方法STELLSH已经作为STELLLS程序的一部分实施:http://www.engr.uconn.edu/sellywu/STELLS.html。联系人:ywu@engr.uconn.eduSupplementary information:补充数据可在Bioinformaticsonline获得。
Motivation:Population trees represent past population divergence histories. The inference of population trees can be useful for the study of population evolution. With the size of data increases in large-scale population genetic projects, such as the 1000 Genomes Project, there are new computational challenges for ancestral population inference, including population tree inference. Existing methods for population tree inference are mainly designed for unlinked genetic variants (e.g. single nucleotide polymorphisms or SNPs). There is a potential loss of information by not considering the haplotypes.Results:In this article, we propose a new population tree inference method (called STELLSH) based on coalescent likelihood. The likelihood is for haplotypes over multiple SNPs within a non-recombining region, not unlinked variants. Unlike many existing ancestral inference methods, STELLSHdoes not use Monte Carlo approaches when computing the likelihood. For efficient computation, the likelihood model is approximated but still retains much information about population divergence history. STELLSHcan find the maximum likelihood population tree based on the approximate likelihood. We show through simulation data and the 1000 Genomes Project data that STELLSHgives reasonably accurate inference results. STELLSHis reasonably efficient for data of current interest and can scale to handle whole-genome data.Availability and implementation:The population tree inference method STELLSHhas been implemented as part of the STELLS program: http://www.engr.uconn.edu/∼ywu/STELLS.html.Contact:ywu@engr.uconn.eduSupplementary information:Supplementary Data are available atBioinformaticsonline.