Improved BLAST searches using longer words for protein seeding

Improved BLAST searches using longer words for protein seeding
复制标题

DOI:
10.1093/bioinformatics/btm479
复制
发表时间:
2007-11-01
期刊:
影响因子:
5.8
通讯作者:
Agarwala, Richa
Agarwala, Richa
中科院分区:
生物学3区
文献类型:
--
作者:
Shiryev, Sergey A.;Papadopoulos, Jason S.;Agarwala, Richa

文献摘要

被引文献

相似文献

动机:BLAST的blastp和tblastp模块是分别针对蛋白质和核苷酸数据库搜索蛋白质查询的广泛使用的方法。在BLAST中使用的一种启发式方法是仅考虑包含与查询的长度至多为5的高分匹配的数据库序列。我们实现了使用长度为6或7的单词的功能。我们证明了一个改进的运行时间和检索精度之间的权衡,控制用于短词匹配的分数阈值。例如,运行时间可以减少20-30,同时实现与使用当前默认参数获得的ROC(受试者操作员特征)评分相似的ROC评分。
Motivation: The blastp and tblastn modules of BLAST are widely used methods for searching protein queries against protein and nucleotide databases, respectively. One heuristic used in BLAST is to consider only database sequences that contain a high-scoring match of length at most 5 to the query. We implemented the capability to use words of length 6 or 7. We demonstrate an improved trade-off between running time and retrieval accuracy, controlled by the score threshold used for short word matches. For example, the running time can be reduced by 20-30 while achieving ROC (receiver operator characteristic) scores similar to those obtained with current default parameters.