Measurement of the effectiveness of transitive sequence comparison, through a third 'intermediate' sequence

Measurement of the effectiveness of transitive sequence comparison, through a third 'intermediate' sequence
复制标题

DOI:
10.1093/bioinformatics/14.8.707
复制
发表时间:
1998-01-01
期刊:
影响因子:
5.8
通讯作者:
Gerstein, M
Gerstein, M
中科院分区:
生物学3区
文献类型:
--
作者:
Gerstein, M

文献摘要

被引文献

相似文献

动机:传递序列匹配通过将数据库中给定查询的结果作为新查询重新运行来扩展序列比较的范围。有时,这会导致初始查询序列 (Q) 通过第三个“中间”序列 (Q --> I --> M) 间接与最终匹配 (M) 相关。这种方法经常被认为可以在序列比较中提供更高的灵敏度;然而;目前还不可能精确地衡量其改进。结果:在这里,通过查看传递序列匹配可以揭示超出正常成对比较(即直接连锁)发现的已知结构关系的哪些部分来全面测量这种改进。结构关系取自一个充分表征的测试集,即蛋白质结构的范围分类。具体来说,2055 个关系较远的蛋白质之间已知的结构相似性(称为“对”)构成了基本测试集。为了正确测量传递匹配,由此衍生出称为“基线集”的特殊数据集。它们由具有清晰结构关系的序列对组成,无法通过正常的序列比较关系找到。它们不能直接链接)。具体来说,使用标准序列比较协议(FASTA,e 值截止值为 0.001),发现基线集由 1742 对组成。第三个中间序列可以间接连接其中的 86 个(5%),其中该第三个序列取自整个当前的蛋白质序列。误报的数量很少。此外,当仅考虑测试集中与紧密结构对齐相对应的关系时,覆盖范围会大大增加。特别是 862 个基线集对的拟合度优于 2.6 埃 RMS,传递匹配可以找到其中的 62 个 (9%)。
Motivation: Transitive sequence matching expands the scope of sequence comparison by re-running the results of a given query against the databank as a new query. This sometimes results in the initial query sequence (Q) being related to a final match (M) indirectly, through a third, 'intermediate' sequence (Q --> I --> M). This approach has often been suggested as providing greater sensitivity in sequence comparison; however; it has not yet been possible to gauge its improvement precisely.Results: Here, this improvement is comprehensively measured by seeing what fraction of the known structural relationships transitive sequence matching can uncover beyond that found by normal pairwise comparison (i.e. direct linkage). The structural relationships are taken from a well-characterized test set, the scop classification of protein structure. Specifically, 2055 known structural similarities (called 'pairs') between distantly related proteins constitute the basic test set. To make the measurement of transitive matching properly, special data sets, called 'baseline sets', are derived from this. They consist of pairs of sequences that have a clear structural relationship that cannot be found by normal sequence comparison tie. they cannot be directly linked). Specifically, using standard sequence comparison protocols (FASTA with an e-value cut-off of 0.001), it is found that the baseline set consists of 1742 pairs. A third intermediate sequence can link 86 of these indirectly (5%), where this third sequence is drawn from the entire, current universe of protein sequences. The number of false positives is minimal. Furthermore, when one considers only the relationships within the test set that correspond to a close structural alignment, the coverage increases considerably. In particular 862 of the baseline set pairs fit to better than 2.6 Angstrom RMS, and transitive matching can find 62 of these (9%).