Optimizing substitution matrices by separating score distributions

Optimizing substitution matrices by separating score distributions
复制标题

DOI:
10.1093/bioinformatics/btg494
复制
发表时间:
2004-04-12
期刊:
影响因子:
5.8
通讯作者:
Akiyama, Y
Akiyama, Y
中科院分区:
生物学3区
文献类型:
--
作者:
Hourai, Y;Akutsu, T;Akiyama, Y

文献摘要

被引文献

相似文献

动机:同源搜索是生物信息学中最基本的工具之一。典型的比对算法使用替换矩阵和缺口代价。因此,替换矩阵的改进提高了同源搜索的准确性。通常,替换矩阵是从已知其关系的比对序列中得出的,并且缺口成本由试错法确定。利用贝叶斯决策理论,从正反两个角度对替换矩阵进行优化,以更清楚地区分关系。结果:利用COG数据库,优化了替换矩阵。得到的矩阵对COG数据库的分类精度优于传统的替换矩阵。在与其他数据库的分类中也取得了较好的效果。
Motivation:Homology search is one of the most fundamental tools in Bioinformatics. Typical alignment algorithms use substitution matrices and gap costs. Thus, the improvement of substitution matrices increases accuracy of homology searches. Generally, substitution matrices are derived from aligned sequences whose relationships are known, and gap costs are determined by trial and error. To discriminate relationships more clearly, we are encouraged to optimize the substitution matrices from statistical viewpoints using both positive and negative examples utilizing Bayesian decision theory.Results: Using Cluster of Orthologous Group (COG) database, we optimized substitution matrices. The classification accuracy of the obtained matrix is better than that of conventional substitution matrices to COG database. It also achieves good performance in classifying with other databases.