Statistical evaluation of pairwise protein sequence comparison with the Bayesian bootstrap

Statistical evaluation of pairwise protein sequence comparison with the Bayesian bootstrap
复制标题

DOI:
10.1093/bioinformatics/bti627
复制
发表时间:
2005-10-15
期刊:
影响因子:
5.8
通讯作者:
Brenner, SE
Brenner, SE
中科院分区:
生物学3区
文献类型:
--
作者:
Price, GA;Crooks, GE;Brenner, SE

文献摘要

被引文献

相似文献

动机:蛋白质序列比较方法通常用于推断在快速增长的蛋白质序列库中发现的复杂的进化关系网络,从而预测未表征蛋白质的结构和功能。在本研究中,我们详细介绍了一种改进的统计基准的成对蛋白质序列比较算法。我们使用自举响应技术来确定标准统计误差,并估计我们的结论的置信度。我们发现,基准数据库内的底层结构导致埃夫隆的标准,非参数引导是有偏见的。因此,标准引导低估了平均性能时,用于评估序列比较方法的背景下。我们已经开发出,作为一种替代方案,一个无偏的统计评价贝叶斯引导,rescovery方法操作类似于标准bootstrap.Results:我们应用我们的分析的氨基酸取代矩阵家庭的比较研究,发现使用现代矩阵结果在一个小的,但统计学上显着改善远程同源性检测相比,经典的PAM和BLOSUM矩阵。
Motivation: Protein sequence comparison methods are routinely used to infer the intricate network of evolutionary relationships found within the rapidly growing library of protein sequences, and thereby to predict the structure and function of uncharacterized proteins. In the present study, we detail an improved statistical benchmark of pairwise protein sequence comparison algorithms. We use bootstrap resampling techniques to determine standard statistical errors and to estimate the confidence of our conclusions. We show that the underlying structure within benchmark databases causes Efron's standard, non-parametric bootstrap to be biased. Consequently, the standard bootstrap underpredicts average performance when used in the context of evaluating sequence comparison methods. We have developed, as an alternative, an unbiased statistical evaluation based on the Bayesian bootstrap, a resampling method operationally similar to the standard bootstrap.Results: We apply our analysis to the comparative study of amino acid substitution matrix families and find that using modern matrices results in a small, but statistically significant improvement in remote homology detection compared with the classic PAM and BLOSUM matrices.