Taxon sampling and the accuracy of large phylogenies
Taxon sampling and the accuracy of large phylogenies
复制标题
DOI:
10.1080/106351598260680
复制
发表时间:
1998-12-01
影响因子:
6.5
通讯作者:
Nielsen, R
中科院分区:
文献类型:
--
作者:
Rannala, B;Huelsenbeck, JP;Nielsen, R
With the advent of automated methods for rapid sequencing of DNA, and the avail? ability of powerful microcomputers, many more attempts are being made to recon? struct large phylogenetic trees that may in? clude hundreds of sequences and thousands of sites (Vigilant et al, 1991; Chase et al., 1993; Krings et al., 1997). This technologi? cal revolution has forced systematists to fo? cus renewed attention on issues relating to the effects of sampling on the accuracy of reconstructed phylogenies. One avenue of research investigates the effect that charac? ter sampling has on phylogenetic estima? tion. That is, the number of taxa in the anal? ysis is held constant but different samples of characters are drawn to investigate the accuracy of phylogenetic methods for sam? ples comprising different numbers of sites and different genomic regions (Graybeal, 1994; Cummings et al, 1995). Another av? enue of research has investigated the effect of taxon sampling on phylogenetic accuracy. For example, Hendy and Penny (1989) ex? amined the consistency of the maximum parsimony (MP) method of inferring phy? logeny for cases in which the molecular clock assumption (that substitution rates do not vary among lineages) is satisfied. They studied a simple (Poisson process) model of substitution, with only two possible states for each character, and focused on five and six taxon trees. They found that the longest branches were attracted to one another in the MP tree; the MP method can therefore be inconsistent (ie, the estimated phylogeny will converge to an incorrect phylogeny as the number of independent characters in the analysis is increased), even in cases where the rates of substitution are equal among lin? eages and arbitrarily low. Hendy and Penny suggested that judicious addition of taxa can break up long branches and help the MP method to become consistent. Kim (1996) tested this prediction, using a combination of analytic theory and computer simulation, and argued that the problem of inconsis? tency becomes worse for the MP method as the number of taxa in the analysis increases. However, the manner in which the taxa were added to the analysis was somewhat unreal? istic; instead of adding taxa within a mono? phyletic group, as most systematists attempt to do, Kim (1996) increased the age of the root while adding more taxa. Most recently, Hillis (1996) has studied the effect of taxon sampling on phyloge? netic accuracy more directly by attempting to evaluate the accuracy of a reconstructed phylogeny for a real taxonomic group, the angiosperms, which includes large numbers of taxa. The phylogeny of 228 species of angiosperms was first inferred from com? plete 18S ribosomal RNA genes by use of the MP method. Artificial data sets were then generated by computer simulation, based on the estimated phylogeny and a model of nucleotide substitution, after which the accuracy of phylogenies estimated from the artificial data sets, using either MP or neighbor-joining (NJ) methods, was deter? mined. A remarkable result was that both procedures appear able to accurately recon? struct the phylogeny for this large group of taxa, using DNA sequences of only a few thousand sites. This is in sharp con? trast with earlier simulation results, and em? pirical studies, which have suggested that much larger sequences may not provide suf? ficient information to accurately estimate phylogeny for as few as four taxa (Hillis et al, 1994).