Calculating bootstrap probabilities of phylogeny using multilocus sequence data

Calculating bootstrap probabilities of phylogeny using multilocus sequence data
复制标题

DOI:
10.1093/molbev/msn043
复制
发表时间:
2008-05-01
影响因子:
10.7
通讯作者:
Seo, Tae-Kun
Seo, Tae-Kun
中科院分区:
生物学1区
文献类型:
--
作者:
Seo, Tae-Kun

文献摘要

被引文献

相似文献

系统发育估计在分子进化研究中起着至关重要的作用。可用基因组数据量的增加有助于从多位点序列数据进行系统发育估计。尽管极大似然和贝叶斯方法可用于多位点序列数据的系统发育重建,但这些方法需要大量的计算,并且它们的应用仅限于中等数量的基因和分类群的分析。距离矩阵方法为分析大量序列数据提供了合适的选择。然而,距离方法应用于多位点序列数据的方式仍然未知。在此,我们提出了利用多位点序列数据估计分子系统发育的新方法,并评估其在距离法框架下的意义。我们发现,多位点序列数据的串联可能会导致错误的系统发育估计,并且具有极高的自举概率(BP),这是由于对距离的错误估计和故意忽略基因间变异。因此,我们建议单独估计多位点序列数据的距离矩阵,然后将这些矩阵组合起来重建系统发育,而不是使用串联序列数据进行系统发育重建。为了计算重建系统发育的bp值,我们建议采用两阶段bootstrap方法;在这种情况下,基因被重新采样,然后在重新采样的基因序列列的重新采样。在计算bp时,通过重新采样基因,适当地考虑了基因间变异。通过仿真研究和经验数据分析,我们证明了我们的两阶段自举过程比序列拼接后采用的传统自举过程更适合。
Phylogeny estimation is extremely crucial in the study of molecular evolution. The increase in the amount of available genomic data facilitates phylogeny estimation from multilocus sequence data. Although maximum likelihood and Bayesian methods are available for phylogeny reconstruction using multilocus sequence data, these methods require heavy computation, and their application is limited to the analysis of a moderate number of genes and taxa. Distance matrix methods present suitable alternatives for analyzing huge amounts of sequence data. However, the manner in which distance methods can be applied to multilocus sequence data remains unknown. Here, we suggest new procedures to estimate molecular phylogeny using multilocus sequence data and evaluate its significance in the framework of the distance method. We found that concatenation of the multilocus sequence data may result in incorrect phylogeny estimation with an extremely high bootstrap probability (BP), which is due to incorrect estimation of the distances and intentional ignorance of the intergene variations. Therefore, we suggest that the distance matrices for multilocus sequence data be estimated separately and these matrices be subsequently combined to reconstruct phylogeny instead of phylogeny reconstruction using concatenated sequence data. To calculate the BPs of the reconstructed phylogeny, we suggest that 2-stage bootstrap procedures be adopted; in this, genes are resampled followed by resampling of the sequence columns within the resampled genes. By resampling the genes during calculation of BPs, intergene variations are properly considered. Via simulation studies and empirical data analysis, we demonstrate that our 2-stage bootstrap procedures are more suitable than the conventional bootstrap procedure that is adopted after sequence concatenation.