Improved global protein homolog detection with major gains in function identification.
Improved global protein homolog detection with major gains in function identification.
复制标题
DOI:
10.1073/pnas.2211823120
复制
发表时间:
2023-02-28
影响因子:
11.1
通讯作者:
中科院分区:
文献类型:
--
作者:
Homolog detection, finding similar proteins to an unknown protein, is usually the first step in understanding the role and function of that protein. However, if the identity of protein sequences between query and target proteins is low (< 30%), traditional tools struggle to distinguish a correct match from a random one, failing to identify important similarities. We have used protein representations from deep learning language models to solve this problem. Reducing the size of these representations significantly improved homolog detection capabilities. Our tool can find putative homologs for more than 93% of human proteins that were not able to assign a function as of March 2022. There are several hundred million protein sequences, but the relationships among them are not fully available from existing homolog detection methods. There is an essential need for an improved method to push homolog detection to lower levels of sequence identity. The method used here relies on a language model to represent proteins numerically in a matrix (an embedding) and uses discrete cosine transforms to compress the data to extract the most essential part, significantly reducing the data size. This PRotein Ortholog Search Tool (PROST) is significantly faster with linear runtimes, and most importantly, computes the distances between pairs of protein sequences to yield homologs at significantly lower levels of sequence identity than previously. The extent of allosteric effects in proteins points out the importance of global aspects of structure and sequence. PROST excels at global homology detection but not at detecting local homologs. Results are validated by strong similarities between the corresponding pairs of structures. The number of remote homologs detected increased significantly and pushes the effective sequence matches more deeply into the twilight zone. Human protein sequences presently having no assigned function now find significant numbers of putative homologs for 93% of cases and structurally verified assigned functions for 76.4% of these cases. The data compression enables massive searches for homologs with short search times while yielding significant gains in the numbers of remote homologs detected. The method is sufficiently efficient to permit whole-genome/proteome comparisons. The PROST web server is accessible at https://mesihk.github.io/prost.
登录
查看更多内容
影响因子:
14.9
作者:
Lees JG;Lee D;Studer RA;Dawson NL;Sillitoe I;Das S;Yeats C;Dessailly BH;Rentzsch R;Orengo CA
通讯作者:
Orengo CA
影响因子:
64.8
作者:
Jumper J;Evans R;Pritzel A;Green T;Figurnov M;Ronneberger O;Tunyasuvunakool K;Bates R;Žídek A;Potapenko A;Bridgland A;Meyer C;Kohl SAA;Ballard AJ;Cowie A;Romera-Paredes B;Nikolov S;Jain R;Adler J;Back T;Petersen S;Reiman D;Clancy E;Zielinski M;Steinegger M;Pacholska M;Berghammer T;Bodenstein S;Silver D;Vinyals O;Senior AW;Kavukcuoglu K;Kohli P;Hassabis D
通讯作者:
Hassabis D
影响因子:
14.9
作者:
UniProt Consortium
通讯作者:
UniProt Consortium
影响因子:
5.7
作者:
Ingles-Prieto, Alvaro;Ibarra-Molero, Beatriz;Delgado-Delgado, Asuncion;Perez-Jimenez, Raul;Fernandez, Julio M.;Gaucher, Eric A.;Sanchez-Ruiz, Jose M.;Gavira, Jose A.
通讯作者:
Gavira, Jose A.
影响因子:
5.6
作者:
Blake, JD;Cohen, FE
通讯作者:
Cohen, FE