Learning Dictionaries for Information Extraction by Multi-Level Bootstrapping

Learning Dictionaries for Information Extraction by Multi-Level Bootstrapping
复制标题

DOI:
--
复制
发表时间:
1999-07
期刊:
--
影响因子:
--
通讯作者:
E. Riloff;R. Jones
E. Riloff;R. Jones
中科院分区:
其他
文献类型:
--
作者:
E. Riloff;R. Jones

文献摘要

被引文献

相似文献

信息抽取系统通常需要两个词典:一个语义词典和一个领域抽取模式词典。我们提出了一种同时生成语义词典和抽取模式的多级自举算法。作为输入,我们的技术只需要未注释的训练文本和少数类别的种子单词。我们使用互引导技术交替地为类别选择最优的抽取模式,并将其抽取的内容引导到语义词典中,这是选择下一个抽取模式的基础。为了使这种方法更健壮,我们添加了第二级引导(元引导),该级别只保留由相互引导产生的最可靠的词典条目,然后重新启动该过程。我们在一组公司网页和一组恐怖主义新闻文章语料库上评估了这种多级自举技术。该算法为几个语义类别生成了高质量的词典。
Information extraction systems usually require two dictionaries: a semantic lexicon and a dictionary of extraction patterns for the domain. We present a multilevel bootstrapping algorithm that generates both the semantic lexicon and extraction patterns simultaneously. As input, our technique requires only unannotated training texts and a handful of seed words for a category. We use a mutual bootstrapping technique to alternately select the best extraction pattern for the category and bootstrap its extractions into the semantic lexicon, which is the basis for selecting the next extraction pattern. To make this approach more robust, we add a second level of bootstrapping (metabootstrapping) that retains only the most reliable lexicon entries produced by mutual bootstrapping and then restarts the process. We evaluated this multilevel bootstrapping technique on a collection of corporate web pages and a corpus of terrorism news articles. The algorithm produced high-quality dictionaries for several semantic categories.