Natural language processing in text mining for structural modeling of protein complexes.

Natural language processing in text mining for structural modeling of protein complexes.
复制标题

DOI:
10.1186/s12859-018-2079-4
复制
发表时间:
2018-03-05
期刊:
影响因子:
3
通讯作者:
Vakser IA
Vakser IA
中科院分区:
生物学4区
文献类型:
--
作者:
Badal VD;Kundrotas PJ;Vakser IA

文献摘要

参考文献

被引文献

相似文献

蛋白质-蛋白质相互作用的结构模型产生了大量蛋白质复合物的假定构型。识别其中的近乎原生模型是一个严峻的挑战。公开的生物医学研究结果可能会对结合模式提供限制,这对于对接至关重要。我们的文本挖掘 (TM) 工具可从 PubMed 摘要中提取结合位点残基,已成功应用于蛋白质对接(Badal 等人,PLoS Comput Biol,2015;11:e1004630)。尽管如此,许多提取的残留物与对接无关。我们提出了 TM 工具的扩展,它利用自然语言处理(NLP)来分析残基出现的上下文。该过程使用通用和专用词典进行了测试。结果表明,为识别蛋白质相互作用而设计的关键词词典不足以用于结合模式的TM预测。然而,我们的字典旨在区分与蛋白质结合位点相关的关键词,从而显着提高了 TM 性能。我们基于对句子解析树的剖析,研究了几种上下文分析方法的实用性。基于机器学习的 NLP 比基于规则的 NLP 更有效地过滤挖掘的残基池。 NLP 生成的约束在来自 DOCKGROUND X 射线基准集 4 的未结合蛋白质的对接中进行了测试。全局低分辨率对接扫描的输出分别通过基本 TM 的约束、NLP 重新排序的约束以及参考约束进行后处理。匹配的质量通过界面均方根偏差来评估。结果表明,当使用高级 TM 与 NLP 生成的约束时,对接输出显着提高。通过深度解析(用于上下文分析的 NLP 技术)清除提取的残基的初始池,显着改进了从 PubMed 摘要中提取蛋白质-蛋白质结合位点残基的基本 TM 程序。基准测试显示,基于高级 TM 与 NLP 产生的约束,对接成功率大幅提高。本文的在线版本 (10.1186/s12859-018-2079-4) 包含补充材料,可供授权用户使用。
Structural modeling of protein-protein interactions produces a large number of putative configurations of the protein complexes. Identification of the near-native models among them is a serious challenge. Publicly available results of biomedical research may provide constraints on the binding mode, which can be essential for the docking. Our text-mining (TM) tool, which extracts binding site residues from the PubMed abstracts, was successfully applied to protein docking (Badal et al., PLoS Comput Biol, 2015; 11: e1004630). Still, many extracted residues were not relevant to the docking. We present an extension of the TM tool, which utilizes natural language processing (NLP) for analyzing the context of the residue occurrence. The procedure was tested using generic and specialized dictionaries. The results showed that the keyword dictionaries designed for identification of protein interactions are not adequate for the TM prediction of the binding mode. However, our dictionary designed to distinguish keywords relevant to the protein binding sites led to considerable improvement in the TM performance. We investigated the utility of several methods of context analysis, based on dissection of the sentence parse trees. The machine learning-based NLP filtered the pool of the mined residues significantly more efficiently than the rule-based NLP. Constraints generated by NLP were tested in docking of unbound proteins from the DOCKGROUND X-ray benchmark set 4. The output of the global low-resolution docking scan was post-processed, separately, by constraints from the basic TM, constraints re-ranked by NLP, and the reference constraints. The quality of a match was assessed by the interface root-mean-square deviation. The results showed significant improvement of the docking output when using the constraints generated by the advanced TM with NLP. The basic TM procedure for extracting protein-protein binding site residues from the PubMed abstracts was significantly advanced by the deep parsing (NLP techniques for contextual analysis) in purging of the initial pool of the extracted residues. Benchmarking showed a substantial increase of the docking success rate based on the constraints generated by the advanced TM with NLP. The online version of this article (10.1186/s12859-018-2079-4) contains supplementary material, which is available to authorized users.
DOI: 10.1371/journal.pcbi.1004630
发表时间: 2015-12
影响因子: 4.3
作者:
Badal VD;Kundrotas PJ;Vakser IA
通讯作者: Vakser IA
DOI: 10.1109/mis.2002.999215
发表时间: 2002-03-01
影响因子: 6.4
作者:
Blaschke, C;Valencia, A
通讯作者: Valencia, A
DOI: 10.1136/jamia.1994.95236146
发表时间: 1994-03-01
影响因子: 6.4
作者:
FRIEDMAN, C;ALDERSON, PO;JOHNSON, SB
通讯作者: JOHNSON, SB
DOI: 10.1155/2015/928531
发表时间: 2015
影响因子: --
作者:
Koyabu S;Phan TT;Ohkawa T
通讯作者: Ohkawa T
DOI: 10.1093/bioinformatics/btl616
发表时间: 2007-02-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Fundel, Katrin;Kueffner, Robert;Zimmer, Ralf
通讯作者: Zimmer, Ralf