A SEQUENCE PROPERTY APPROACH TO SEARCHING PROTEIN DATABASES

A SEQUENCE PROPERTY APPROACH TO SEARCHING PROTEIN DATABASES
复制标题

DOI:
10.1006/jmbi.1995.0442
复制
发表时间:
1995-08-18
影响因子:
5.6
通讯作者:
SANDER, C
SANDER, C
中科院分区:
生物学2区
文献类型:
--
作者:
HOBOHM, U;SANDER, C

文献摘要

被引文献

相似文献

目前可用的序列比对程序通常不能检测序列相似性的过渡区中的功能和结构同源物,即当序列同一性福尔斯低于约25%时。在这里,我们试图检测这种弱的相似性使用的方法的基础上的概念,蛋白质序列相似性从根本上不同的顺序比对。该方法定义了蛋白质序列相异性氨基酸组成(例如,单线态和双线态氨基酸组成、分子量、等电点)的差异的加权和使用PropSearch,单个序列可以用于数据库查询,或者可以将多个序列合并成反映蛋白质家族的平均组成的“平均"序列。首先,我们表明,结构蛋白质家族的成员有一个低的相互PropSearch距离时的权重进行了优化,以最大限度地区分结构家族。其次,我们展示了使用PropSearch方法进行数据库搜索的结果。当扫描预处理的数据库时,这种搜索非常快速,并且不需要比对。在常规比对工具无法检测相似性的情况下,PropSearch可以用于生成关于新序列与数据库中序列之间可能的结构或功能关系的假设。
Currently available sequence alignment programs are generally not capable of detecting functional and structural homologs in the twilight zone of sequence similarity, i.e. when the sequence identity falls below about 25%. Here we attempt to detect such weak similarities using an approach based on a notion of protein sequence similarity radically different from that used in sequential alignment. The approach defines protein sequence dissimilarity (or distance) as a weighted sum of differences of compositional properties such as singlet and doublet amino acid composition, molecular weight, isoelectric point (protein property search or PropSearch).With PropSearch, either single sequences can be used for a database query, or multiple sequences can be merged into an ''average'' sequence reflecting the average composition of a protein family.First, we show that members of structural protein families have a low mutual PropSearch distance when the weights are optimized to discriminate maximally between structural families. Second, we demonstrate the results of database searches using the PropSearch method. Such searches are very rapid when scanning a preprocessed database and do not require alignments.In cases in which conventional alignment tools fail to detect similarities, PropSearch can be used to generate hypotheses about possible structural or functional relationships between a new sequence and sequences in the database.