Detecting remote evolutionary relationships among proteins by large-scale semantic embedding.
Detecting remote evolutionary relationships among proteins by large-scale semantic embedding.
复制标题
DOI:
10.1371/journal.pcbi.1001047
复制
发表时间:
2011-01-27
影响因子:
4.3
通讯作者:
Leslie C
中科院分区:
文献类型:
--
作者:
Melvin I;Weston J;Noble WS;Leslie C
Virtually every molecular biologist has searched a protein or DNA sequence database to find sequences that are evolutionarily related to a given query. Pairwise sequence comparison methods—i.e., measures of similarity between query and target sequences—provide the engine for sequence database search and have been the subject of 30 years of computational research. For the difficult problem of detecting remote evolutionary relationships between protein sequences, the most successful pairwise comparison methods involve building local models (e.g., profile hidden Markov models) of protein sequences. However, recent work in massive data domains like web search and natural language processing demonstrate the advantage of exploiting the global structure of the data space. Motivated by this work, we present a large-scale algorithm called ProtEmbed, which learns an embedding of protein sequences into a low-dimensional “semantic space.” Evolutionarily related proteins are embedded in close proximity, and additional pieces of evidence, such as 3D structural similarity or class labels, can be incorporated into the learning process. We find that ProtEmbed achieves superior accuracy to widely used pairwise sequence methods like PSI-BLAST and HHSearch for remote homology detection; it also outperforms our previous RankProp algorithm, which incorporates global structure in the form of a protein similarity network. Finally, the ProtEmbed embedding space can be visualized, both at the global level and local to a given query, yielding intuition about the structure of protein sequence space. Searching a protein or DNA sequence database to find sequences that are evolutionarily related to a query is one of the foundational problems in computational biology. These database searches rely on pairwise comparisons of sequence similarity between the query and targets, but despite years of method refinements, pairwise comparisons still often fail to detect more distantly related targets. In this study, we adapt recent work from natural language processing to exploit the global structure of the data space in this detection problem. In particular, we borrow the idea of a semantic embedding, where by training on a large text data set, one learns an embedding of words into a low-dimensional semantic space such that words embedded close to each other are likely to be semantically related. We present the ProtEmbed algorithm, which learns an embedding of protein sequences into a semantic space where evolutionarily-related proteins are embedded in close proximity. The flexible training algorithm allows additional pieces of evidence, such as 3D structural information, to be incorporated in the learning process and enables ProtEmbed to achieve state-of-the-art performance for the task of detecting targets that have remote evolutionary relationships to the query.
登录
查看更多内容
DOI:
10.1073/pnas.0308067101
发表时间:
2004-04-27
影响因子:
11.1
作者:
Weston, J;Elisseeff, A;Noble, WS
通讯作者:
Noble, WS
DOI:
10.1111/1467-9868.00346
发表时间:
2002-01-01
影响因子:
5.8
作者:
Storey, JD
通讯作者:
Storey, JD
DOI:
10.1111/j.2517-6161.1995.tb02031.x
发表时间:
1995-01-01
影响因子:
5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者:
HOCHBERG, Y
影响因子:
14.9
作者:
Söding, J;Biegert, A;Lupas, AN
通讯作者:
Lupas, AN
影响因子:
8
作者:
Rychlewski, L;Jaroszewski, L;Godzik, A
通讯作者:
Godzik, A