Constructing Biological Knowledge Bases by Extracting Information from Text Sources

Constructing Biological Knowledge Bases by Extracting Information from Text Sources
复制标题

DOI:
--
复制
发表时间:
1999-08
期刊:
Proceedings. International Conference on Intelligent Systems for Molecular Biology
影响因子:
--
通讯作者:
M. Craven;J. Kumlien
M. Craven;J. Kumlien
中科院分区:
其他
文献类型:
--
作者:
M. Craven;J. Kumlien

文献摘要

被引文献

相似文献

最近,人们在使分子生物学数据库更容易访问和互操作方面做了很多努力。然而,文本形式的信息,如MEDLINE记录,仍然是一个大大未被利用的生物信息来源。我们已经开始了一项研究工作,旨在自动映射信息从文本源到结构化的表示,如知识库。我们完成这项任务的方法是使用机器学习方法来归纳从文本中提取事实的例程。我们描述了应用于此任务的两种学习方法--统计文本分类方法和关系学习方法--以及我们在学习此类信息提取例程方面的初步实验。我们还提出了一种方法,通过从“弱”标记的训练数据中学习来降低学习信息提取例程的成本。
Recently, there has been much effort in making databases for molecular biology more accessible and interoperable. However, information in text form, such as MEDLINE records, remains a greatly underutilized source of biological information. We have begun a research effort aimed at automatically mapping information from text sources into structured representations, such as knowledge bases. Our approach to this task is to use machine-learning methods to induce routines for extracting facts from text. We describe two learning methods that we have applied to this task--a statistical text classification method, and a relational learning method--and our initial experiments in learning such information-extraction routines. We also present an approach to decreasing the cost of learning information-extraction routines by learning from "weakly" labeled training data.