Bootstrap Confidence Levels for Phylogenetic Trees

Bootstrap Confidence Levels for Phylogenetic Trees
复制标题

系统发育树的 Bootstrap 置信水平

DOI:
10.1007/978-0-387-75692-9_17
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
J. Felsenstein
J. Felsenstein
中科院分区:
--
文献类型:
--
作者:
J. Felsenstein

文献摘要

被引文献

相似文献

在20世纪80年代初,生物学家对重建进化树(英语:Approximatenies)越来越感兴趣,通常使用DNA序列。越来越清楚的是,这一点应在统计上加以处理。但是,尽管点估计可以通过最大似然法或最小二乘法进行,但可能的树的空间足够大,而且足够奇怪,因此不清楚如何知道对树的特征有多大的信心。在1985年,我建议Efron的非参数自助法可以应用于这个问题(Felsenstein 1985)。如果我们有一个数据表,其中行是不同的物种,列是DNA分子中的不同位点,我们将通过绘制列(保持每列中的行顺序相同)进行自举,直到我们有一个具有相同物种和相同位点数的数据集。对于每一个自举样本数据集,我们都可以推断出一棵树,一个重要的问题是树的一个给定的分支(例如导致人类和黑猩猩而不是大猩猩或猩猩的分支)是否可以被验证。我天真地认为,分支的存在与否是一个离散的0/1变量,并假设我们可以使用百分位数方法和自助法来确定这个变量的分布质量的95%是否在一个原子中。如果这个分支出现在从自举样本推断出的树中超过95%,我们应该声明它是显著支持的。该方法满足了需求,并被广泛使用-Ryan和Woodall(2005)将其列为引用最多的统计论文列表中的第7位。我很惊讶地发现,它的排名超过了埃夫隆关于自举的原始论文,直到我意识到这份名单并不是统计学家的引用。因此,它倾向于强调应用,而不是作为应用基础的理论。埃夫隆的论文介绍了自举是更有影响力的,但有很多人推断出遗传学。
In the early 1980s, biologists were increasingly interested in reconstructing phylogenies (evolutionary trees), often using DNA sequences. It was becoming clear that this should be treated statistically. But although point estimates could be made by maximum likelihood or least squares methods, the space of possible trees was large enough, and strange enough, that it was not clear how to know how much confidence to have in features of the tree. In 1985 I suggested that Efron’s nonparametric bootstrap could be applied to the problem (Felsenstein 1985). If we have a table of data with rows being different species and columns being different sites in the DNA molecule, we would bootstrap by drawing columns (keeping the rows in the same order within each column), until we had a data set with the same species and the same number of sites. For each of these bootstrap sample data sets, we would infer a tree.An important question was whether a given branch of the tree (such as the branch that leads to humans and chimpanzees but not to gorillas or orangutans) could be validated. I naïvely thought of the presence or absence of the branch as a discrete 0/1 variable and assumed that we could use the percentile method with the bootstrap to see whether (say) 95% of the mass of the distribution of this variable was in the one atom. We should declare the branch significantly supported if it appeared in more than 95% of the trees inferred from bootstrap samples. The method met a need and was very widely used–Ryan and Woodall (2005) list it as number 7 in a list of most-cited statistical papers. I was astonished to see that it outranked Efron’s original paper on the bootstrap, until I realized that the list was not of citations by statisticians. It thus tended to emphasize applications rather than the theory underlying them. Efron’s paper introducing the bootstrap was far more influential, but there were a great many people inferring phylogenies.