The predictive power of the CluSTr database

The predictive power of the CluSTr database
复制标题

DOI:
10.1093/bioinformatics/bti542
复制
发表时间:
2005-09-15
期刊:
影响因子:
5.8
通讯作者:
Apweiler, R
Apweiler, R
中科院分区:
生物学3区
文献类型:
--
作者:
Petryszak, R;Kretschmann, E;Apweiler, R

文献摘要

被引文献

相似文献

CLUSTER数据库采用基于相似性矩阵的全自动单链分层聚类方法。为了计算矩阵,使用Smith-Waterman算法计算蛋白质序列之间的首先全合理比较。然后,使用蒙特卡洛分析评估相似性得分的统计显着性,从而产生z值,这些Z值用于填充基质。本文介绍了自动注释实验,以量化预测能力,从而量化clustrust数据的生物学相关性。该实验利用Uniprot数据挖掘框架来使用Interpro和Clustrust的组合来得出注释预测。我们表明,与仅使用Interpro数据相比,这种数据源的组合大大提高了数据挖掘框架进行的预测精度。我们得出的结论是,聚集蛋白的固定方法对传统蛋白质分类做出了宝贵的贡献。
The CluSTr database employs a fully automatic single-linkage hierarchical clustering method based on a similarity matrix. In order to compute the matrix, first all-against-all pair-wise comparisons between protein sequences are computed using the Smith-Waterman algorithm. The statistical significance of the similarity scores is then assessed using a Monte Carlo analysis, yielding Z-values, which are used to populate the matrix. This paper describes automated annotation experiments that quantify the predictive power and hence the biological relevance of the CluSTr data. The experiments utilized the UniProt data-mining framework to derive annotation predictions using combinations of InterPro and CluSTr. We show that this combination of data sources greatly increases the precision of predictions made by the data-mining framework, compared with the use of InterPro data alone. We conclude that the CluSTr approach to clustering proteins makes a valuable contribution to traditional protein classifications.