Text mining improves prediction of protein functional sites.

Text mining improves prediction of protein functional sites.
复制标题

DOI:
10.1371/journal.pone.0032171
复制
发表时间:
2012
期刊:
影响因子:
3.7
通讯作者:
Wall ME
Wall ME
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Verspoor KM;Cohn JD;Ravikumar KE;Wall ME

文献摘要

参考文献

被引文献

相似文献

我们提出了一种方法,集成蛋白质结构分析和文本挖掘的蛋白质功能位点预测,称为LEAP-FS(文献增强自动预测功能位点)。使用动力学扰动分析(DPA)进行结构分析,该分析预测相互作用极大地扰动蛋白质振动的控制点处的功能位点。文本挖掘提取文献中提到的残留物,并预测所提到的残留物在功能上是重要的。我们通过分析这些方法在大约100,000个公开可用的蛋白质结构中发现已知功能位点(特别是小分子结合位点和催化位点)的性能,评估了每种方法的重要性。DPA预测概括了许多功能位点注释,并且优先恢复注释为生物学相关的结合位点与注释为潜在虚假的结合位点。基于文本的预测也得到了功能位点注释的实质性支持:与其他残基相比,文本中提到的残基在功能位点中被发现的可能性大约是其他残基的六倍。当基于文本和基于结构的方法一致时,预测与注释的重叠得到改善。我们的分析还产生了许多功能位点残基的新的高质量预测,这些残基没有在我们检查的数据源中编目。我们的结论是,DPA和文本挖掘独立地提供了有价值的高通量蛋白质功能位点预测,并集成使用LEAP-FS这两种方法进一步提高了这些预测的质量。
We present an approach that integrates protein structure analysis and text mining for protein functional site prediction, called LEAP-FS (Literature Enhanced Automated Prediction of Functional Sites). The structure analysis was carried out using Dynamics Perturbation Analysis (DPA), which predicts functional sites at control points where interactions greatly perturb protein vibrations. The text mining extracts mentions of residues in the literature, and predicts that residues mentioned are functionally important. We assessed the significance of each of these methods by analyzing their performance in finding known functional sites (specifically, small-molecule binding sites and catalytic sites) in about 100,000 publicly available protein structures. The DPA predictions recapitulated many of the functional site annotations and preferentially recovered binding sites annotated as biologically relevant vs. those annotated as potentially spurious. The text-based predictions were also substantially supported by the functional site annotations: compared to other residues, residues mentioned in text were roughly six times more likely to be found in a functional site. The overlap of predictions with annotations improved when the text-based and structure-based methods agreed. Our analysis also yielded new high-quality predictions of many functional site residues that were not catalogued in the curated data sources we inspected. We conclude that both DPA and text mining independently provide valuable high-throughput protein functional site predictions, and that integrating the two methods using LEAP-FS further improves the quality of these predictions.
DOI: 10.1007/s10796-006-6103-2
发表时间: 2006-02-01
影响因子: 5.9
作者:
Baker, CJO;Witte, R
通讯作者: Witte, R
DOI: 10.1093/bioinformatics/18.9.1280
发表时间: 2002-09-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Greer, DS;Westbrook, JD;Bourne, PE
通讯作者: Bourne, PE
DOI: 10.1107/s0907444904033189
发表时间: 2005-03-01
影响因子: 2.2
作者:
Huang, LL;Zhao, XM;Xia, ZX
通讯作者: Xia, ZX
DOI: 10.1021/bi000574g
发表时间: 2000-07-25
期刊: BIOCHEMISTRY
影响因子: 2.9
作者:
Choe, JY;Fromm, HJ;Honzatko, RB
通讯作者: Honzatko, RB
DOI: 10.1016/s1359-0278(97)00024-2
发表时间: 1997-01-01
期刊: FOLDING & DESIGN
影响因子: --
作者:
Bahar, I;Atilgan, AR;Erman, B
通讯作者: Erman, B