Reducing Labeling Effort for Structured Prediction Tasks

Reducing Labeling Effort for Structured Prediction Tasks
复制标题

DOI:
10.21236/ada440382
复制
发表时间:
2005-07
期刊:
影响因子:
2.1
通讯作者:
A. Culotta;A. McCallum
A. Culotta;A. McCallum
中科院分区:
材料科学3区
文献类型:
--
作者:
A. Culotta;A. McCallum

文献摘要

被引文献

相似文献

阻碍监督机器学习算法快速部署的一个常见障碍是缺乏标记的训练数据。对于结构化的预测任务来说,这是非常昂贵的,因为每个训练实例可能有多个相互作用的标签,所有这些标签都必须被正确地注释,以使实例对学习者有用。传统的主动学习通过优化标记示例的顺序来解决这个问题,以提高学习效率。然而,这种方法没有考虑标记每个示例的难度,这在结构化预测任务中变化很大。例如,由部分训练的系统预测的标签可能在某些情况下比在其他情况下更容易纠正。我们提出了一种新的主动学习范式,它不仅减少了注释者必须标记的实例数量,而且减少了每个实例的注释难度。该系统还利用来自部分正确预测的信息来有效地征求用户的注释。我们在交互式信息提取系统中验证了这种主动学习框架,将注释操作总数减少了22%。
A common obstacle preventing the rapid deployment of supervised machine learning algorithms is the lack of labeled training data. This is particularly expensive to obtain for structured prediction tasks, where each training instance may have multiple, interacting labels, all of which must be correctly annotated for the instance to be of use to the learner. Traditional active learning addresses this problem by optimizing the order in which the examples are labeled to increase learning efficiency. However, this approach does not consider the difficulty of labeling each example, which can vary widely in structured prediction tasks. For example, the labeling predicted by a partially trained system may be easier to correct for some instances than for others. We propose a new active learning paradigm which reduces not only how many instances the annotator must label, but also how difficult each instance is to annotate. The system also leverages information from partially correct predictions to efficiently solicit annotations from the user. We validate this active learning framework in an interactive information extraction system, reducing the total number of annotation actions by 22%.