Comparative experiments on learning information extractors for proteins and their interactions

Comparative experiments on learning information extractors for proteins and their interactions
复制标题

DOI:
10.1016/j.artmed.2004.07.016
复制
发表时间:
2005-02-01
影响因子:
7.5
通讯作者:
Wong, YW
Wong, YW
中科院分区:
工程技术1区
文献类型:
--
作者:
Bunescu, R;Ge, RF;Wong, YW

文献摘要

被引文献

相似文献

目的:从生物医学文本中自动提取信息有望以计算机可访问的形式轻松整合大量生物学知识。这种策略对于从Medline中的1100万摘要中提取与人类基因组基因相关的数据特别有吸引力。然而,由于缺乏描述人类基因和蛋白质的惯例,提取工作一直受到挫折。我们已经开发和评估了各种学习信息提取系统,用于识别Medtine摘要中的人类蛋白质名称,并随后提取蛋白质之间相互作用的信息。方法和材料:我们使用了各种机器学习方法来自动开发信息提取系统,用于从Medline摘要中提取基因/蛋白质名称,功能和相互作用的信息。我们提出了交叉验证的结果,通过训练和测试一组约1000手动注释的Medline摘要,讨论人类基因/蛋白质识别人类蛋白质及其相互作用。结果如下:我们证明了使用支持向量机和最大熵的机器学习方法能够比以前的几种方法更准确地识别人类蛋白质。我们还表明,各种规则诱导方法能够识别蛋白质相互作用,具有更高的精度比手动开发的rules.Conclusion:我们的研究结果表明,它是有前途的,使用机器学习自动构建系统,从生物医学文本中提取信息。结果还提供了一个广泛的图片的相对优势,各种各样的方法进行测试时,一个合理的大型人类注释的语料库。(c)2004 Elsevier B. V.保留所有权利。
Objective: Automatically extracting information from biomedical text holds the promise of easily consolidating large amounts of biological knowledge in computer-accessible form. This strategy is particularly attractive for extracting data relevant to genes of the human genome from the 11 million abstracts in Medline. However, extraction efforts have been frustrated by the lack of conventions for describing human genes and proteins. We have developed and evaluated a variety of learned information extraction systems for identifying human protein names in Medtine abstracts and subsequently extracting information on interactions between the proteins.Methods and Material: We used a variety of machine learning methods to automaticatly develop information extraction systems for extracting information on gene/ protein name, function and interactions from Medline abstracts. We present crossvalidated results on identifying human proteins and their interactions by training and testing on a set of approximately 1000 manuatly-annotated Medline abstracts that discuss human genes/proteins. Results: We demonstrate that machine learning approaches using support vector machines and maximum entropy are able to identify human proteins with higher accuracy than several previous approaches. We also demonstrate that various rule induction methods are able to identify protein interactions with higher precision than manually-developed rules.Conclusion: Our results show that it is promising to use machine learning to automatically build systems for extracting information from biomedical text. The results also give a broad picture of the relative strengths of a wide variety of methods when tested on a reasonably large human-annotated corpus. (c) 2004 Elsevier B.V. All rights reserved.