Detecting remote evolutionary relationships among proteins by large-scale semantic embedding.

Detecting remote evolutionary relationships among proteins by large-scale semantic embedding.
复制标题

DOI:
10.1371/journal.pcbi.1001047
复制
发表时间:
2011-01-27
影响因子:
4.3
通讯作者:
Leslie C
Leslie C
中科院分区:
生物学2区
文献类型:
--
作者:
Melvin I;Weston J;Noble WS;Leslie C

文献摘要

参考文献

被引文献

相似文献

几乎每个分子生物学家都搜索过蛋白质或DNA序列数据库,以寻找与给定查询在进化上相关的序列。成对序列比较方法--即查询和目标序列之间的相似性度量--为序列数据库搜索提供了引擎,并已成为30年来计算研究的主题。对于检测蛋白质序列之间的远程进化关系这一难题,最成功的两两比较方法包括建立蛋白质序列的局部模型(例如,轮廓隐马尔可夫模型)。然而,最近在网络搜索和自然语言处理等海量数据领域的研究表明,利用数据空间的全局结构具有优势。受这项工作的启发,我们提出了一个名为ProtEmbed的大规模算法,该算法学习将蛋白质序列嵌入到低维“语义空间”中。进化上相关的蛋白质嵌入在非常接近的位置,额外的证据,如3D结构相似性或类别标签,可以被纳入学习过程。我们发现,与广泛使用的PSI-BLAST和HHSearch等用于远程同源性检测的成对序列方法相比,ProtEmbed具有更高的准确性;它也优于我们之前的RankProp算法,后者以蛋白质相似性网络的形式结合了全局结构。最后,ProtEmbedding嵌入空间可以在全局级别和给定查询的局部级别上可视化,从而直观地了解蛋白质序列空间的结构。搜索蛋白质或DNA序列数据库以找到与查询在进化上相关的序列是计算生物学的基本问题之一。这些数据库搜索依赖于查询和目标之间的序列相似性的成对比较,但尽管多年的方法改进,成对比较仍然经常无法检测到更远的相关目标。在本研究中,我们采用了自然语言处理的最新工作来探索这个检测问题中数据空间的全局结构。特别是,我们借用了语义嵌入的想法,其中通过在大型文本数据集上的训练,人们学习了将单词嵌入到低维语义空间中,从而使彼此嵌入的单词可能在语义上相关。我们提出了ProtEmed算法,它学习将蛋白质序列嵌入到语义空间中,在语义空间中,进化相关的蛋白质被嵌入到紧密接近的位置。灵活的训练算法允许将其他证据片段(如3D结构信息)合并到学习过程中,并使ProtEmed能够在检测与查询具有远程进化关系的目标的任务中实现最先进的性能。
Virtually every molecular biologist has searched a protein or DNA sequence database to find sequences that are evolutionarily related to a given query. Pairwise sequence comparison methods—i.e., measures of similarity between query and target sequences—provide the engine for sequence database search and have been the subject of 30 years of computational research. For the difficult problem of detecting remote evolutionary relationships between protein sequences, the most successful pairwise comparison methods involve building local models (e.g., profile hidden Markov models) of protein sequences. However, recent work in massive data domains like web search and natural language processing demonstrate the advantage of exploiting the global structure of the data space. Motivated by this work, we present a large-scale algorithm called ProtEmbed, which learns an embedding of protein sequences into a low-dimensional “semantic space.” Evolutionarily related proteins are embedded in close proximity, and additional pieces of evidence, such as 3D structural similarity or class labels, can be incorporated into the learning process. We find that ProtEmbed achieves superior accuracy to widely used pairwise sequence methods like PSI-BLAST and HHSearch for remote homology detection; it also outperforms our previous RankProp algorithm, which incorporates global structure in the form of a protein similarity network. Finally, the ProtEmbed embedding space can be visualized, both at the global level and local to a given query, yielding intuition about the structure of protein sequence space. Searching a protein or DNA sequence database to find sequences that are evolutionarily related to a query is one of the foundational problems in computational biology. These database searches rely on pairwise comparisons of sequence similarity between the query and targets, but despite years of method refinements, pairwise comparisons still often fail to detect more distantly related targets. In this study, we adapt recent work from natural language processing to exploit the global structure of the data space in this detection problem. In particular, we borrow the idea of a semantic embedding, where by training on a large text data set, one learns an embedding of words into a low-dimensional semantic space such that words embedded close to each other are likely to be semantically related. We present the ProtEmbed algorithm, which learns an embedding of protein sequences into a semantic space where evolutionarily-related proteins are embedded in close proximity. The flexible training algorithm allows additional pieces of evidence, such as 3D structural information, to be incorporated in the learning process and enables ProtEmbed to achieve state-of-the-art performance for the task of detecting targets that have remote evolutionary relationships to the query.
DOI: 10.1073/pnas.0308067101
发表时间: 2004-04-27
影响因子: 11.1
作者:
Weston, J;Elisseeff, A;Noble, WS
通讯作者: Noble, WS
DOI: 10.1111/1467-9868.00346
发表时间: 2002-01-01
影响因子: 5.8
作者:
Storey, JD
通讯作者: Storey, JD
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
用于蛋白质同源性检测和结构预测的 HHpred 交互式服务器。
DOI: 10.1093/nar/gki408
发表时间: 2005-07-01
影响因子: 14.9
作者:
Söding, J;Biegert, A;Lupas, AN
通讯作者: Lupas, AN
DOI: 10.1110/ps.9.2.232
发表时间: 2000-02-01
期刊: PROTEIN SCIENCE
影响因子: 8
作者:
Rychlewski, L;Jaroszewski, L;Godzik, A
通讯作者: Godzik, A