Sentence-Based Active Learning Strategies for Information Extraction

Sentence-Based Active Learning Strategies for Information Extraction
复制标题

基于句子的主动学习策略用于信息提取

DOI:
--
复制
发表时间:
2010
期刊:
Italian Information Retrieval Workshop
影响因子:
--
通讯作者:
F. Sebastiani
F. Sebastiani
中科院分区:
--
文献类型:
--
作者:
Andrea Esuli;Diego Marcheggiani;F. Sebastiani

文献摘要

被引文献

相似文献

给定在相对较少的训练示例上训练的分类器,主动学习(AL)包括根据它们的信息量对一组未标记的示例进行排序,如果手动标记,则可以重新训练(希望)更好的分类器。信息提取(IE)是一个重要的文本学习任务,在这个任务中,人工智能可能是有用的,即在文本中识别实例化给定概念的表达式的任务。我们认为,与其他文本学习任务不同,IE的独特之处在于,对单个项目(即单词出现次数)进行注释排序没有意义,并且呈现给注释者的最小文本单元应该是一个完整的句子。在本文中,我们提出了一系列基于单个句子排序的IE主动学习策略,并在命名实体提取的标准数据集上对它们进行了实验比较。
Given a classifier trained on relatively few training examples, active learning (AL) consists in ranking a set of unlabeled examples in terms of how informative they would be, if manually labeled, for retraining a (hopefully) better classifier. An important text learning task in which AL is potentially useful is information extraction (IE), namely, the task of identifying within a text the expressions that instantiate a given concept. We contend that, unlike in other text learning tasks, IE is unique in that it does not make sense to rank individual items (i.e., word occurrences) for annotation, and that the minimal unit of text that is presented to the annotator should be an entire sentence. In this paper we propose a range of active learning strategies for IE that are based on ranking individual sentences, and experimentally compare them on a standard dataset for named entity extraction.