A Local Alignment Metric for Accelerating Biosequence Database Search

A Local Alignment Metric for Accelerating Biosequence Database Search
复制标题

用于加速生物序列数据库搜索的局部比对指标

DOI:
10.1089/106652704773416894
复制
发表时间:
2004
期刊:
Journal of computational biology : a journal of computational molecular cell biology
影响因子:
--
通讯作者:
Natasa Macura
Natasa Macura
中科院分区:
--
文献类型:
--
作者:
Peter A. Spiro;Natasa Macura

文献摘要

参考文献

被引文献

相似文献

我们引入了一种用于局部序列比对的度量,该度量可用于加速最佳比对搜索而不损失灵敏度。该度量的三角不等式属性允许识别冗余数据库条目,保证与低于指定分数阈值的查询序列具有最佳对齐,从而允许跳过与这些条目的比较。我们证明了各种评分系统(包括最常用的评分系统)的度量的存在,并表明对于核苷酸与蛋白质序列的比较也可以建立三角不等式。我们讨论利用三角不等式的数据库聚类和搜索策略。该策略允许对广泛使用的“nr”蛋白质数据库进行适度但显着的搜索加速。它还提供了一种基于理论的数据库聚类方法,并提供了比较启发式聚类策略的标准。
We introduce a metric for local sequence alignments that has utility for accelerating optimal alignment searches without loss of sensitivity. The metric's triangle inequality property permits identification of redundant database entries guaranteed to have optimal alignments to the query sequence that fall below a specified score threshold, thereby permitting comparisons to these entries to be skipped. We prove the existence of the metric for a variety of scoring systems, including the most commonly used ones, and show that a triangle inequality can be established as well for nucleotide-to-protein sequence comparisons. We discuss a database clustering and search strategy that takes advantage of the triangle inequality. The strategy permits moderate but significant acceleration of searches against the widely used "nr" protein database. It also provides a theoretically based method for database clustering in general and provides a standard against which to compare heuristic clustering strategies.
使用比对分数的上限对数据库序列进行聚类,以进行快速同源搜索。
DOI: --
发表时间: 2004
期刊: Genome Informatics 15
影响因子: --
作者:
Itoh;M.;Akutsu;T.;Kanehisa;M.
通讯作者: M.