Identifying GPCR-drug interaction based on wordbook learning from sequences

Identifying GPCR-drug interaction based on wordbook learning from sequences
复制标题

基于从序列中学习的单词书识别 GPCR-药物相互作用

DOI:
10.1186/s12859-020-3488-8
复制
发表时间:
2020-04
期刊:
影响因子:
3
通讯作者:
Xiao Xuan
Xiao Xuan
中科院分区:
生物学4区
文献类型:
--
作者:
Wang Pu;Huang Xiaotong;Qiu Wangren;Xiao Xuan

文献摘要

参考文献

相似文献

背景G蛋白偶联受体(GPCRs)介导多种重要的生理功能,与多种疾病密切相关,是现代药物最重要的靶向家族。因此,GPCR分析和GPCR配体筛选的研究是新药开发的热点。准确识别GPCR靶向药物的作用是设计GPCR靶向药物的关键步骤之一。然而,在实验上大规模确定GPCR-药物对的相互作用的成本高得令人望而却步。因此,直接从分子序列预测GPCR-药物对的相互作用具有重要意义。结果提出了一种新的基于序列的GPCR-药物相互作用识别方法。对于GPCRS,我们使用了一种新颖的词袋模型(BOW)来提取序列特征,该模型可以从低阶到高阶提取更多的模式信息,并限制了特征空间维度。对于药物分子,我们使用离散傅立叶变换(DFT)从原始的分子指纹中提取高阶模式信息。将两种分子的特征向量连接起来,输入到一个简单的预测引擎--距离加权K近邻(DWKNN)。这种基本方法很容易通过集成学习来增强。通过对最近构建的GPCR-药物相互作用数据集的测试,发现所提出的方法在泛化能力上优于现有的基于序列的机器学习方法,甚至是一种通过后处理过程(PPP)进一步提高预测性能的非常规方法。结论所提出的方法对于GPCR-药物相互作用的预测是有效的,也可能成为其他靶向-药物相互作用预测或蛋白质-蛋白质相互作用预测的潜在方法。此外,新提出的GPCR序列特征提取方法是传统BOW模型的改进版本,可能有助于解决蛋白质分类或属性预测问题。建议方法的源代码可在https://github.com/wp3751/GPCR-Drug-Interaction.上免费获得,供学术研究使用
BackgroundG protein-coupled receptors (GPCRs) mediate a variety of important physiological functions, are closely related to many diseases, and constitute the most important target family of modern drugs. Therefore, the research of GPCR analysis and GPCR ligand screening is the hotspot of new drug development. Accurately identifying the GPCR-drug interaction is one of the key steps for designing GPCR-targeted drugs. However, it is prohibitively expensive to experimentally ascertain the interaction of GPCR-drug pairs on a large scale. Therefore, it is of great significance to predict the interaction of GPCR-drug pairs directly from the molecular sequences. With the accumulation of known GPCR-drug interaction data, it is feasible to develop sequence-based machine learning models for query GPCR-drug pairs.ResultsIn this paper, a new sequence-based method is proposed to identify GPCR-drug interactions. For GPCRs, we use a novel bag-of-words (BoW) model to extract sequence features, which can extract more pattern information from low-order to high-order and limit the feature space dimension. For drug molecules, we use discrete Fourier transform (DFT) to extract higher-order pattern information from the original molecular fingerprints. The feature vectors of two kinds of molecules are concatenated and input into a simple prediction engine distance-weighted K-nearest-neighbor (DWKNN). This basic method is easy to be enhanced through ensemble learning. Through testing on recently constructed GPCR-drug interaction datasets, it is found that the proposed methods are better than the existing sequence-based machine learning methods in generalization ability, even an unconventional method in which the prediction performance was further improved by post-processing procedure (PPP).ConclusionsThe proposed methods are effective for GPCR-drug interaction prediction, and may also be potential methods for other target-drug interaction prediction, or protein-protein interaction prediction. In addition, the new proposed feature extraction method for GPCR sequences is the modified version of the traditional BoW model and may be useful to solve problems of protein classification or attribute prediction. The source code of the proposed methods is freely available for academic research at https://github.com/wp3751/GPCR-Drug-Interaction.
DOI: 10.1093/bioinformatics/btn162
发表时间: 2008-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Yamanishi Y;Araki M;Gutteridge A;Honda W;Kanehisa M
通讯作者: Kanehisa M
DOI: 10.1093/nar/27.1.368
发表时间: 1999-01-01
影响因子: 14.9
作者:
Kawashima, S;Ogata, H;Kanehisa, M
通讯作者: Kanehisa, M
iGPCR-Drug:用于预测蜂窝网络中 GPCR 和药物之间相互作用的 Web 服务器
DOI: 10.1371/journal.pone.0072234
发表时间: 2013
期刊: PloS one
影响因子: 3.7
作者:
Xiao X;Min JL;Wang P;Chou KC
通讯作者: Chou KC
DOI: 10.1038/s41598-019-43125-6
发表时间: 2019-05-22
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
Li, Li;Koh, Ching Chiek;Wei, Dong-Qing
通讯作者: Wei, Dong-Qing