TCS: a computer program to estimate gene genealogies

TCS: a computer program to estimate gene genealogies
复制标题

DOI:
10.1046/j.1365-294x.2000.01020.x
复制
发表时间:
2000-10-01
期刊:
影响因子:
4.9
通讯作者:
Crandall, KA
Crandall, KA
中科院分区:
生物学1区
文献类型:
--
作者:
Clement, M;Posada, D;Crandall, KA

文献摘要

被引文献

相似文献

系统发育是非常有用的工具,不仅用于建立一组生物或其部分(如基因)之间的系谱关系,而且一旦估计了系统发育,还可用于各种研究。在最近的综述中,Pagel(1999)雄辩地概述了系统发育信息的许多用途,从发现耐药性到重建所有生命的共同祖先。系统发育已被用于预测传染病的未来趋势(Bush et al. 1999),甚至被用作法庭证据(Vogel 1997)。然而,遗传学的有用之处在于它们的准确性,在群体水平上估计基因之间的系谱关系给传统的遗传重建方法带来了许多困难。这些传统的方法,如简约法,邻域连接,最大似然假设是无效的人口水平。例如,这些方法假设祖先单倍型不再存在于群体中,但结合理论预测祖先单倍型将是群体水平研究中最常见的序列(Watterson & Guess 1977; Donnelly & Tavaré 1986; Crandall & Templeton 1993)。传统的方法需要相当大数量的变量字符,以准确地重建关系(Huelsenbeck和希利斯1993)和人口水平的研究通常缺乏这样的变化。此外,在群体水平上,序列之间的重组是真实的可能性,并且传统方法假设重组不发生。如果没有考虑到基因重组的可能性,重建基因组可能会导致严重的错误,在所产生的估计基因组。这些效应的组合可以导致简约方法在群体水平上推断出大量最简约的树,而在集合中没有分辨率(例如,对于一组人类线粒体DNA(mtDNA),超过10亿棵树,Excoffier & Smouse 1994)。这些影响也会导致相邻连接和传统的最大似然方法对结果关系过于自信(Bandelt et al. 1995)。因此,需要一种替代方法来提供对群体基因谱系的准确估计
Phylogenies are extremely useful tools, not only for establishing genealogical relationships among a group of organisms or their parts (eg genes), but also for a variety of research once the phylogenies are estimated. In a recent review, Pagel (1999) eloquently outline a number of uses for phylogenetic information from discovery of drug resistance to reconstructing the common ancestor to all of life. Phylogenies have been used to predict future trends in infectious disease (Bush et al. 1999) and have even been offered as evidence in a court of law (Vogel 1997). Yet phylogenies are only as useful as they are accurate.Estimating genealogical relationships among genes at the population level presents a number of difficulties to traditional methods of phylogeny reconstruction. These traditional methods such as parsimony, neighbour-joining, and maximumlikelihood make assumptions that are invalid at the population level. For example, these methods assume ancestral haplotypes are no longer in the population, yet coalescent theory predicts that ancestral haplotypes will be the most frequent sequences sampled in a population level study (Watterson & Guess 1977; Donnelly & Tavaré 1986; Crandall & Templeton 1993). Traditional methods require reasonably large numbers of variable characters to accurately reconstruct relationships (Huelsenbeck & Hillis 1993) and population level studies typically lack such variation. Also, recombination is a real possibility among sequences at the population level and traditional methods assume recombination does not occur. The failure to incorporate the possibility of recombination in phylogeny reconstruction can lead to grave errors in the resulting estimated phylogeny. The combination of these effects can lead parsimony methods to infer a cumbersome amount of most parsimonious trees at the population level with no resolution among the set (eg over one billion trees for a set of human mitochondrial DNA (mtDNA), Excoffier & Smouse 1994). These effects can also lead neighbour-joining and traditional maximum-likelihood methods to be over confident in the resulting relationships (Bandelt et al. 1995). Therefore, an alternative approach is needed to provide accurate estimates of gene genealogies at the population