Strategies for the effective identification of remotely related sequences in multiple PSSM search approach

Strategies for the effective identification of remotely related sequences in multiple PSSM search approach
复制标题

DOI:
10.1002/prot.21356
复制
发表时间:
2007-06-01
影响因子:
2.9
通讯作者:
Srinivasan, N.
Srinivasan, N.
中科院分区:
生物学4区
文献类型:
--
作者:
Gowri, V. S.;Tina, K. G.;Srinivasan, N.

文献摘要

被引文献

相似文献

使用位置特定评分矩阵(PSSMs)的搜索通常用于远程同源检测程序,如PSI-BLAST和RPS-BLAST。PSSM通常使用一个序列族中的一个序列作为参考序列来生成。在PSI-BLAST搜索的情况下,引用序列与查询相同。最近,我们已经表明,对多个家庭档案数据库的搜索,每个家庭的成员作为参考序列使用,比对单一家庭档案的经典数据库的搜索更有效。尽管与普通序列-概要匹配过程相比,它的总体性能相对较好,但针对多个家族概要数据库的搜索会导致一些假阳性和假阴性。本文表明,构建PSSM时所使用的序列的轮廓长度和散度对基于多轮廓的搜索方法的性能有重要影响。我们还发现,一个简单的参数定义为一个查询所命中的家族对应的pssm的数量,除以家族中pssm的总数,可以有效地区分多配置文件搜索方法中的真阳性和假阳性。蛋白质2007;67:789 - 794。(C) 2007 Wiley-Liss, Inc。
Searches using position specific scoring matrices (PSSMs) have been commonly used in remote homology detection procedures such as PSI-BLAST and RPS-BLAST. A PSSM is generated typically using one of the sequences of a family as the reference sequence. In the case of PSI-BLAST searches the reference sequence is same as the query. Recently we have shown that searches against the database of multiple family-profiles, with each one of the members of the family used as a reference sequence, are more effective than searches against the classical database of single family-profiles. Despite relatively a better overall performance when compared with common sequence-profile matching procedures, searches against the multiple family-profiles database result in a few false positives and false negatives. Here we show that profile length and divergence of sequences used in the construction of a PSSM have major influence on the performance of multiple profile based search approach. We also identify that a simple parameter defined by the number of PSSMs corresponding to a family that is hit, for a query, divided by the total number of PSSMs in the family can distinguish effectively the true positives from the false positives in the multiple profiles search approach. Proteins 2007;67:789-794. (C) 2007 Wiley-Liss, Inc.