Text mining for modeling of protein complexes enhanced by machine learning

Text mining for modeling of protein complexes enhanced by machine learning
复制标题

DOI:
10.1093/bioinformatics/btaa823
复制
发表时间:
2021-02-15
期刊:
影响因子:
5.8
通讯作者:
Vakser,Ilya A.
Vakser,Ilya A.
中科院分区:
生物学3区
文献类型:
--
作者:
Badal,Varsha D.;Kundrotas,Petras J.;Vakser,Ilya A.

文献摘要

被引文献

相似文献

蛋白质-蛋白质复合物(蛋白质对接)的结构建模过程产生了许多需要进一步分析和评分的模型。评分可以基于对复合物结构的独立确定的约束,例如对蛋白质相互作用所必需的氨基酸的了解。在此之前,我们发现在免费的PubMed关于蛋白质-蛋白质相互作用研究的论文摘要中对残基进行文本挖掘可能会产生这样的限制。然而,缺乏对斑点残基的后处理降低了约束的可用性,因为大量残基与特定蛋白质的结合无关。结果我们通过深度递归神经网络(DRNN)和支持向量机(SVM)两种机器学习方法,采用不同的训练/测试方案,探索了不相关残基的过滤。结果表明,当对PMC-OA全文文章进行训练并将其应用于PubMed摘要中发现的残基分类(界面或非界面)时,DRNN模型优于SVM模型。当对全文文章或摘要进行训练和测试时,这些模型的性能是相似的。因此,在这种情况下,不需要使用计算要求很高的DRNN方法,这种方法在计算上非常昂贵,特别是在训练阶段。原因是支持向量机的成功与否往往取决于训练集和测试集中数据/文本模式的相似性,而摘要中的句子结构通常与全文文章中的句子结构不同。可获得性和实施本研究生成的代码和数据集可在https://gitlab.ku.edu/vakser-lab-public/text-mining/-/tree/2020-09-04.Supplementary information上获得,补充数据可在bioinformatics online上获得。
MotivationProcedures for structural modeling of protein–protein complexes (protein docking) produce a number of models which need to be further analyzed and scored. Scoring can be based on independently determined constraints on the structure of the complex, such as knowledge of amino acids essential for the protein interaction. Previously, we showed that text mining of residues in freely available PubMed abstracts of papers on studies of protein–protein interactions may generate such constraints. However, absence of post-processing of the spotted residues reduced usability of the constraints, as a significant number of the residues were not relevant for the binding of the specific proteins.ResultsWe explored filtering of the irrelevant residues by two machine learning approaches, Deep Recursive Neural Network (DRNN) and Support Vector Machine (SVM) models with different training/testing schemes. The results showed that the DRNN model is superior to the SVM model when training is performed on the PMC-OA full-text articles and applied to classification (interface or non-interface) of the residues spotted in the PubMed abstracts. When both training and testing is performed on full-text articles or on abstracts, the performance of these models is similar. Thus, in such cases, there is no need to utilize computationally demanding DRNN approach, which is computationally expensive especially at the training stage. The reason is that SVM success is often determined by the similarity in data/text patterns in the training and the testing sets, whereas the sentence structures in the abstracts are, in general, different from those in the full text articles.Availabilityand implementationThe code and the datasets generated in this study are available at https://gitlab.ku.edu/vakser-lab-public/text-mining/-/tree/2020-09-04.Supplementary informationSupplementary data are available atBioinformaticsonline.