YAKUSA: A fast structural database scanning method

YAKUSA: A fast structural database scanning method
复制标题

DOI:
10.1002/prot.20517
复制
发表时间:
2005-10-01
影响因子:
2.9
通讯作者:
Pothier, J
Pothier, J
中科院分区:
生物学4区
文献类型:
--
作者:
Carpentier, M;Brouillet, S;Pothier, J

文献摘要

被引文献

相似文献

YAKUSA是一个程序,设计用于快速扫描结构数据库与查询蛋白质结构。它搜索存在于查询结构和结构数据库中的每个结构之间的最长公共子结构,称为SHSP(结构高分对)。它利用蛋白质骨架内部坐标(α角)来描述蛋白质结构作为符号序列。在5个步骤中建立结构相似性,前3个步骤类似于BLAST中使用的步骤:(1)建立描述与查询结构中的那些相同或相似的所有模式的确定性有限自动机;(2)在数据库中的每个结构中搜索所有这些模式;(3)将模式扩展到更长的匹配子结构(即,(a)小规模社会保障项目;(5)使用基于SHSP相似性、基于SHSP概率和基于SHSP的空间兼容性的3个分数来对查询-数据库结构对进行排名。结构碎片概率根据混合转移分布模型估计,该模型是高阶马尔可夫链模型的近似。关于结构匹配的灵敏度和选择性,YAKUSA与最好的相关程序相比很好,尽管它的速度要快得多:在台式个人计算机上,典型的数据库扫描需要大约40秒的CPU时间。它还在Web服务器上实现了实时搜索。
YAKUSA is a program designed for rapid scanning of a structural database with a query protein structure. It searches for the longest common substructures called SHSPs (structural high-scoring pairs) existing between a query structure and every structure in the structural database. It makes use of protein backbone internal coordinates (alpha angles) in order to describe protein structures as sequences of symbols. The structural similarities are established in 5 steps, the first 3 being analogous to those used in BLAST: (1) building up a deterministic finite automaton describing all patterns identical or similar to those in the query structure; (2) searching for all these patterns in every structure in the database; (3) extending the patterns to longer matching substructures (i.e., SHSPs); (4) selecting compatible SHSPs for each query-database structure pair; and (5) ranking the query- database structure pairs using 3 scores based on SHSP similarity, on SHSP probabilities, and on spatial compatibility of SHSPs. Structural fragment probabilities are estimated according to a mixture transition distribution model, which is an approximation of a high-order Markov chain model. With regard to sensitivity and selectivity of the structural matches, YAKUSA compares well to the best related programs, although it is by far faster: A typical database scan takes about 40 s CPU time on a desktop personal computer. It has also been implemented on a Web server for real-time searches.