On the interpretation of bootstrap trees: Appropriate threshold of clade selection and induced gain

On the interpretation of bootstrap trees: Appropriate threshold of clade selection and induced gain
复制标题

DOI:
10.1093/molbev/13.7.999
复制
发表时间:
1996-09-01
影响因子:
10.7
通讯作者:
Gascuel, O
Gascuel, O
中科院分区:
生物学1区
文献类型:
--
作者:
Berry, V;Gascuel, O

文献摘要

被引文献

相似文献

在这项研究中,我们解决了解释引导树的问题。主要问题是选择进化枝选择的阈值,以便根据其自举比例将可靠的进化枝与不可靠的进化枝分开。该阈值取决于所选的误差测量。我们研究了源自 Robinson 和 Foulds (1981) 距离的概括的误差测量,用于量化真实系统发育和估计树之间的差异。我们提出了进化枝选择最佳阈值的两种分析近似来解释(即减少)引导树。我们按照 Kuhner 和 Felsenstein (1994) 的思路,使用邻接法和最大简约法进行了广泛的模拟。这些模拟表明,与经验观察得出的最佳阈值相比,我们的近似值仅造成很小的质量损失。接下来,我们测量了通过适当减少的引导树而不是通过经典树构建方法获得的完整原始树来估计真实系统发育时所实现的误差减少。我们对短序列的模拟表明,当使用标准 Robinson 和 Foulds 距离测量误差时,简约法可实现 39% 的误差减少,距离法可实现 33% 的误差减少。观察到的错误减少源于 I 类错误(错误的推论)的显着减少,而 II 类错误(忽略正确的进化枝)仅略有增加。当使用较短的序列并且更加重视类型I错误而不是类型II错误时,可以实现更大的错误减少。为了从另一个角度研究误差的原因,我们提出将误差期望一般分解为两项偏差和一项方差。这些项的结果表明,引导过程没有引入基本偏差,唯一的偏差来源是结构性的(缺乏解决方案)。此外,估计的方差大大减少,这为简化的引导树与原始树估计相比的更好结果提供了另一种解释。
In this study we address the problem of interpreting a bootstrap tree. The main issue is choosing the threshold of clade selection in order to separate reliable clades from unreliable ones, depending on their bootstrap proportion. This threshold depends on the chosen error measure. We investigate error measures that stem from a generalization of Robinson and Foulds' (1981) distance, used to quantify the divergence between the true phylogeny and the estimated trees. We propose two analytical approximations of the optimum threshold of clade selection to interpret (i.e., reduce) the bootstrap tree. We performed extensive simulations along the lines of Kuhner and Felsenstein (1994) using the neighbor-joining and the maximum-parsimony methods. These simulations show that our approximations cause only small losses in quality when compared to the optimum threshold resulting from empirical observation. Next, we measured the error reduction achieved when estimating the true phylogeny by the properly reduced bootstrap tree rather than by the complete original tree, obtained with a classical tree-building method. Our simulations on short sequences show that an error reduction of 39% is achieved with the parsimony method and an error reduction of 33% is achieved with the distance method when the error is measured with the standard Robinson and Foulds distance. The observed error reduction is shown to originate from an important decrease in Type I error (wrong inferences), while Type II error (omitted correct clades) is only slightly increased. Greater error reduction is achieved when shorter sequences are used, and when more importance is given to Type I error than to Type II error. To investigate the causes of error from another point of view, we propose a general decomposition of the error expectation in two terms of bias, and one of variance. Results for these terms show that no fundamental bias is introduced by the bootstrap process, the only source of bias being structural (lack of resolution). Moreover, the variance in the estimations is greatly reduced, providing another explanation for the better results of the reduced bootstrap tree compared with the original tree estimate.