Taxon sampling and the accuracy of large phylogenies

Taxon sampling and the accuracy of large phylogenies
复制标题

DOI:
10.1080/106351598260680
复制
发表时间:
1998-12-01
期刊:
影响因子:
6.5
通讯作者:
Nielsen, R
Nielsen, R
中科院分区:
生物学1区
文献类型:
--
作者:
Rannala, B;Huelsenbeck, JP;Nielsen, R

文献摘要

被引文献

相似文献

随着DNA快速测序自动化方法的出现,强大的微型计算机的能力,更多的尝试正在进行侦察?构建大型的系统发育树,包括数百个序列和数千个位点(Vigilant等,1991; Chase等,1993; Krings等人,1997年)。这项技术?卡尔革命迫使系统主义者,cus重新关注与取样对重建遗传学准确性的影响有关的问题。研究的一个途径是调查charac?对系统发育的影响?是的。也就是说,肛门中的分类群数量?ysis是保持不变,但不同的字符样本绘制调查的准确性系统发育方法为山姆?包含不同数量的位点和不同基因组区域的基因组序列(Graybeal,1994; Cummings等,1995)。另一个AV?许多研究调查了分类单元取样对系统发育准确性的影响。例如,亨迪和佩妮(1989)前?证明了最大简约法(MP)推断物理量的一致性。logeny的情况下,分子时钟假设(即取代率不改变血统)得到满足。他们研究了一个简单的(泊松过程)替代模型,每个字符只有两种可能的状态,并专注于五个和六个分类树。他们发现,最长的分支被吸引到另一个MP树; MP方法,因此可以是不一致的(即,估计的同源性将收敛到一个不正确的同源性作为分析中的独立字符的数量增加),即使在案件中的替代率是平等的林?价格和任意低。亨迪和彭尼认为,明智地添加分类群可以打破长分支,并帮助MP方法变得一致。Kim(1996)使用分析理论和计算机模拟相结合的方法检验了这一预测,并认为不一致的问题?随着分析中分类群数量的增加,MP方法的效率变得更差。然而,分类群被添加到分析中的方式有点不真实?而不是在一个单细胞内添加分类群?正如大多数系统分类学家试图做的那样,Kim(1996)增加了根的年龄,同时增加了更多的分类群。最近,希利斯(1996年)研究了分类单元采样对植物的影响?更直接地通过尝试评估一个重建的被子植物系统发生学的准确性,被子植物是一个真实的分类群,它包括大量的分类群。本文首次从玉米属植物中推断出228种被子植物的亲缘关系。用MP法扩增18 S核糖体RNA基因。人工数据集,然后产生的计算机模拟,基于估计的同源性和模型的核苷酸取代,在此之后的准确性估计的同源性从人工数据集,使用MP或相邻连接(NJ)的方法,是deter?地雷。一个值得注意的结果是,这两个程序似乎能够准确地侦察?我们只使用几千个位点的DNA序列,就可以构建这一大群分类群的系统发育。这是骗局吗trast与早期的模拟结果,和EM?peptide的研究,这表明,更大的序列可能不会提供suf?有足够的信息来准确估计少至四个分类群的系统发育(Hillis等,1994)。
With the advent of automated methods for rapid sequencing of DNA, and the avail? ability of powerful microcomputers, many more attempts are being made to recon? struct large phylogenetic trees that may in? clude hundreds of sequences and thousands of sites (Vigilant et al, 1991; Chase et al., 1993; Krings et al., 1997). This technologi? cal revolution has forced systematists to fo? cus renewed attention on issues relating to the effects of sampling on the accuracy of reconstructed phylogenies. One avenue of research investigates the effect that charac? ter sampling has on phylogenetic estima? tion. That is, the number of taxa in the anal? ysis is held constant but different samples of characters are drawn to investigate the accuracy of phylogenetic methods for sam? ples comprising different numbers of sites and different genomic regions (Graybeal, 1994; Cummings et al, 1995). Another av? enue of research has investigated the effect of taxon sampling on phylogenetic accuracy. For example, Hendy and Penny (1989) ex? amined the consistency of the maximum parsimony (MP) method of inferring phy? logeny for cases in which the molecular clock assumption (that substitution rates do not vary among lineages) is satisfied. They studied a simple (Poisson process) model of substitution, with only two possible states for each character, and focused on five and six taxon trees. They found that the longest branches were attracted to one another in the MP tree; the MP method can therefore be inconsistent (ie, the estimated phylogeny will converge to an incorrect phylogeny as the number of independent characters in the analysis is increased), even in cases where the rates of substitution are equal among lin? eages and arbitrarily low. Hendy and Penny suggested that judicious addition of taxa can break up long branches and help the MP method to become consistent. Kim (1996) tested this prediction, using a combination of analytic theory and computer simulation, and argued that the problem of inconsis? tency becomes worse for the MP method as the number of taxa in the analysis increases. However, the manner in which the taxa were added to the analysis was somewhat unreal? istic; instead of adding taxa within a mono? phyletic group, as most systematists attempt to do, Kim (1996) increased the age of the root while adding more taxa. Most recently, Hillis (1996) has studied the effect of taxon sampling on phyloge? netic accuracy more directly by attempting to evaluate the accuracy of a reconstructed phylogeny for a real taxonomic group, the angiosperms, which includes large numbers of taxa. The phylogeny of 228 species of angiosperms was first inferred from com? plete 18S ribosomal RNA genes by use of the MP method. Artificial data sets were then generated by computer simulation, based on the estimated phylogeny and a model of nucleotide substitution, after which the accuracy of phylogenies estimated from the artificial data sets, using either MP or neighbor-joining (NJ) methods, was deter? mined. A remarkable result was that both procedures appear able to accurately recon? struct the phylogeny for this large group of taxa, using DNA sequences of only a few thousand sites. This is in sharp con? trast with earlier simulation results, and em? pirical studies, which have suggested that much larger sequences may not provide suf? ficient information to accurately estimate phylogeny for as few as four taxa (Hillis et al, 1994).