Text Mining for Protein Docking.

Text Mining for Protein Docking.
复制标题

DOI:
10.1371/journal.pcbi.1004630
复制
发表时间:
2015-12
影响因子:
4.3
通讯作者:
Vakser IA
Vakser IA
中科院分区:
生物学2区
文献类型:
--
作者:
Badal VD;Kundrotas PJ;Vakser IA

文献摘要

被引文献

相似文献

来自生物医学研究的公开可用信息量的快速增长在互联网上很容易获得,为预测生物分子建模提供了强大的资源。积累的实验确定的结构的数据转化为蛋白质和蛋白质复合物的结构预测。预测工具可以简单地根据与现有的、先前确定的结构的相似性来进行解决,而不是探索巨大的搜索空间。一个类似的重大范式转变正在出现,由于信息量的迅速扩大,而不是实验确定的结构,这仍然可以用作生物分子结构预测的约束。自动文本挖掘已被广泛用于重建蛋白质相互作用网络,以及在检测蛋白质结构上的小配体结合位点。结合和扩展这两个成熟的研究领域,我们将文本挖掘应用于蛋白质-蛋白质复合物的结构建模(蛋白质对接)。蛋白质对接可以显着改善时,对接模式的限制。我们开发了一个程序,检索发表的摘要上的一个特定的蛋白质-蛋白质相互作用和提取信息对接。在来自Dockground(http://www.example.com)的蛋白质复合物上评估该程序。dockground.compbio.ku.edu结果表明,大约一半的复合物可以提取出正确的结合残基信息。基于词袋(特征)方法,通过对检索到的摘要子集进行概念分析,减少了不相关信息的数量。支持向量机模型的训练和验证的子集。剩余的摘要由性能最好的模型过滤,这减少了数据集中约25%复合物的不相关信息。将提取的约束纳入对接协议中,并在Dockground未绑定基准集上进行测试,显著提高了对接成功率。蛋白质相互作用是许多细胞过程的核心。这些相互作用的物理表征对于理解生命过程以及在生物学和医学中的应用至关重要。由于实验技术的固有局限性以及计算能力和方法论的快速发展,计算机建模是许多研究中的首选工具。来自生物医学研究的公开信息在互联网上很容易获得,为蛋白质和蛋白质复合物的建模提供了强大的资源。蛋白质复合物建模的一个主要范式转变正在出现,由于这些信息的迅速扩大,可用作建模约束。文本挖掘已被广泛用于重建蛋白质相互作用的网络,以及检测蛋白质上的小分子结合位点。结合和扩展这两个成熟的研究领域,我们将文本挖掘应用于蛋白质复合物的物理建模(蛋白质对接)。我们的程序检索发表的摘要蛋白质-蛋白质相互作用,并提取相关信息。结果表明,大约一半的蛋白质复合物可以获得正确的结合信息。提取的约束被纳入建模过程中,显着提高其性能。
The rapidly growing amount of publicly available information from biomedical research is readily accessible on the Internet, providing a powerful resource for predictive biomolecular modeling. The accumulated data on experimentally determined structures transformed structure prediction of proteins and protein complexes. Instead of exploring the enormous search space, predictive tools can simply proceed to the solution based on similarity to the existing, previously determined structures. A similar major paradigm shift is emerging due to the rapidly expanding amount of information, other than experimentally determined structures, which still can be used as constraints in biomolecular structure prediction. Automated text mining has been widely used in recreating protein interaction networks, as well as in detecting small ligand binding sites on protein structures. Combining and expanding these two well-developed areas of research, we applied the text mining to structural modeling of protein-protein complexes (protein docking). Protein docking can be significantly improved when constraints on the docking mode are available. We developed a procedure that retrieves published abstracts on a specific protein-protein interaction and extracts information relevant to docking. The procedure was assessed on protein complexes from Dockground (http://dockground.compbio.ku.edu). The results show that correct information on binding residues can be extracted for about half of the complexes. The amount of irrelevant information was reduced by conceptual analysis of a subset of the retrieved abstracts, based on the bag-of-words (features) approach. Support Vector Machine models were trained and validated on the subset. The remaining abstracts were filtered by the best-performing models, which decreased the irrelevant information for ~ 25% complexes in the dataset. The extracted constraints were incorporated in the docking protocol and tested on the Dockground unbound benchmark set, significantly increasing the docking success rate. Protein interactions are central for many cellular processes. Physical characterization of these interactions is essential for understanding of life processes and applications in biology and medicine. Because of the inherent limitations of experimental techniques and rapid development of computational power and methodology, computer modeling is a tool of choice in many studies. Publicly available information from biomedical research is readily accessible on the Internet, providing a powerful resource for modeling of proteins and protein complexes. A major paradigm shift in modeling of protein complexes is emerging due to the rapidly expanding amount of such information, which can be used as modeling constraints. Text mining has been widely used in recreating networks of protein interactions, as well as in detecting small molecule binding sites on proteins. Combining and expanding these two well-developed areas of research, we applied the text mining to physical modeling of protein complexes (protein docking). Our procedure retrieves published abstracts on a protein-protein interaction and extracts the relevant information. The results show that correct information on binding can be obtained for about half of protein complexes. The extracted constraints were incorporated in a modeling procedure, significantly improving its performance.