CSI: Contrastive data Stratification for Interaction prediction and its application to compound-protein interaction prediction.
CSI: Contrastive data Stratification for Interaction prediction and its application to compound-protein interaction prediction.
复制标题
DOI:
10.1093/bioinformatics/btad456
复制
发表时间:
2023-08-01
期刊:
影响因子:
--
通讯作者:
中科院分区:
文献类型:
--
作者:
Accurately predicting the likelihood of interaction between two objects (compound–protein sequence, user–item, author–paper, etc.) is a fundamental problem in Computer Science. Current deep-learning models rely on learning accurate representations of the interacting objects. Importantly, relationships between the interacting objects, or features of the interaction, offer an opportunity to partition the data to create multi-views of the interacting objects. The resulting congruent and non-congruent views can then be exploited via contrastive learning techniques to learn enhanced representations of the objects. We present a novel method, Contrastive Stratification for Interaction Prediction (CSI), to stratify (partition) a dataset in a manner that can be exploited via Contrastive Multiview Coding to learn embeddings that maximize the mutual information across congruent data views. CSI assigns a key and multiple views to each data point, where data partitions under a particular key form congruent views of the data. We showcase the effectiveness of CSI by applying it to the compound–protein sequence interaction prediction problem, a pressing problem whose solution promises to expedite drug delivery (drug–protein interaction prediction), metabolic engineering, and synthetic biology (compound–enzyme interaction prediction) applications. Comparing CSI with a baseline model that does not utilize data stratification and contrastive learning, and show gains in average precision ranging from 13.7% to 39% using compounds and sequences as keys across multiple drug–target and enzymatic datasets, and gains ranging from 16.9% to 63% using reaction features as keys across enzymatic datasets. Code and dataset available at https://github.com/HassounLab/CSI.
登录
查看更多内容
影响因子:
14.9
作者:
Kanehisa M;Furumichi M;Sato Y;Ishiguro-Watanabe M;Tanabe M
通讯作者:
Tanabe M
DOI:
10.1093/bioinformatics/bty593
发表时间:
2018-09-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
Öztürk H;Özgür A;Ozkirimli E
通讯作者:
Ozkirimli E
影响因子:
9.5
作者:
Bagherian M;Sabeti E;Wang K;Sartor MA;Nikolovska-Coleska Z;Najarian K
通讯作者:
Najarian K
影响因子:
5.8
作者:
Thin Nguyen;Hang Le;Venkatesh, Svetha
通讯作者:
Venkatesh, Svetha
影响因子:
4.3
作者:
Lee, Ingo;Keum, Jongsoo;Nam, Hojung
通讯作者:
Nam, Hojung