PreBIND and Textomy--mining the biomedical literature for protein-protein interactions using a support vector machine.

PreBIND and Textomy--mining the biomedical literature for protein-protein interactions using a support vector machine.
复制标题

DOI:
10.1186/1471-2105-4-11
复制
发表时间:
2003-03-27
期刊:
影响因子:
3
通讯作者:
Hogue, CWV
Hogue, CWV
中科院分区:
生物学4区
文献类型:
--
作者:
Donaldson, I;Martin, J;de Bruijn, B;Wolting, C;Lay, V;Tuekam, B;Zhang, SD;Baskin, B;Bader, GD;Michalickova, K;Pawson, T;Hogue, CWV

文献摘要

参考文献

被引文献

相似文献

大多数实验验证的分子相互作用和生物学途径数据存在于生物医学期刊文章的非结构化文本中,在那里它们无法使用计算方法。生物分子相互作用网络数据库(BIND)试图以机器可读的格式捕获这些数据。我们假设,强大的任务规模的检索数据库可以减少使用支持向量机技术,首先定位在文献中的相互作用信息。我们提出了一个信息提取系统,该系统旨在定位文献中的蛋白质-蛋白质相互作用数据,并将这些数据提供给策展人和公众进行审查并进入BIND。交叉验证估计支持向量机对描述交互信息的摘要分类的测试集精确率、准确率和召回率分别为92%、90%和92%。我们估计,该系统将能够召回高达60%的所有非高通量的相互作用存在于另一个酵母蛋白质相互作用数据库。最后,将该系统应用于现实世界的策展问题,发现其使用可将任务持续时间减少70%,从而节省176天。机器学习方法作为指导交互和路径数据库回填的工具是有用的;然而,只有当这些技术与人类审查和进入事实数据库(如BIND)相结合时,才能实现这种潜力。这里描述的PreBIND系统可供公众使用。目前的能力允许搜索人类,小鼠和酵母蛋白质相互作用的信息。
The majority of experimentally verified molecular interaction and biological pathway data are present in the unstructured text of biomedical journal articles where they are inaccessible to computational methods. The Biomolecular interaction network database (BIND) seeks to capture these data in a machine-readable format. We hypothesized that the formidable task-size of backfilling the database could be reduced by using Support Vector Machine technology to first locate interaction information in the literature. We present an information extraction system that was designed to locate protein-protein interaction data in the literature and present these data to curators and the public for review and entry into BIND. Cross-validation estimated the support vector machine's test-set precision, accuracy and recall for classifying abstracts describing interaction information was 92%, 90% and 92% respectively. We estimated that the system would be able to recall up to 60% of all non-high throughput interactions present in another yeast-protein interaction database. Finally, this system was applied to a real-world curation problem and its use was found to reduce the task duration by 70% thus saving 176 days. Machine learning methods are useful as tools to direct interaction and pathway database back-filling; however, this potential can only be realized if these techniques are coupled with human review and entry into a factual database such as BIND. The PreBIND system described here is available to the public at . Current capabilities allow searching for human, mouse and yeast protein-interaction information.
DOI: 10.1186/1471-2105-3-32
发表时间: 2002-10-25
期刊: BMC bioinformatics
影响因子: 3
作者:
Michalickova K;Bader GD;Dumontier M;Lieu H;Betel D;Isserlin R;Hogue CW
通讯作者: Hogue CW
DOI: 10.1038/88213
发表时间: 2001-05-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Jenssen, TK;Lægreid, A;Hovig, E
通讯作者: Hovig, E
DOI: 10.1093/nar/30.1.31
发表时间: 2002-01-01
影响因子: 14.9
作者:
Mewes, HW;Frishman, D;Weil, B
通讯作者: Weil, B
DOI: 10.1038/nbt1002-991
发表时间: 2002-10-01
影响因子: 46.9
作者:
Bader, GD;Hogue, CWV
通讯作者: Hogue, CWV
DOI: 10.1093/bioinformatics/16.5.465
发表时间: 2000-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bader, GD;Hogue, CWV
通讯作者: Hogue, CWV