Margin-based active learning for structured predictions

Margin-based active learning for structured predictions
复制标题

DOI:
10.1007/s13042-010-0003-y
复制
发表时间:
2010-09
影响因子:
5.6
通讯作者:
Kevin Small;D. Roth
Kevin Small;D. Roth
中科院分区:
计算机科学3区
文献类型:
--
作者:
Kevin Small;D. Roth

文献摘要

被引文献

相似文献

基于边际的主动学习由于其简单性和经验上的成功,仍然是使用最广泛的主动学习范式。然而,大多数工作仅限于二元或多类预测问题,从而限制了这些方法在许多复杂预测问题上的适用性,而主动学习将是最有用的。例如,用于自然语言处理应用的机器学习技术通常需要结合多个相互依赖的预测问题——通常被称为结构化输出空间中的学习。在许多这样的应用程序领域中,通过将复杂的预测分解为一系列预测来进一步管理复杂性,其中早期的预测用作后期预测的输入——通常称为管道模型。这项工作描述了将现有的基于边缘的主动学习技术扩展到这两种设置的方法,从而增加了主动学习可以应用的问题范围。我们通过减少对合成数据的多个实例、语义角色标记任务和命名实体和关系提取系统的注释数据需求,对这些提出的主动学习技术进行了实证验证。
Margin-based active learning remains the most widely used active learning paradigm due to its simplicity and empirical successes. However, most works are limited to binary or multiclass prediction problems, thus restricting the applicability of these approaches to many complex prediction problems where active learning would be most useful. For example, machine learning techniques for natural language processing applications often require combining multiple interdependent prediction problems—generally referred to aslearning in structured output spaces. In many such application domains, complexity is further managed by decomposing a complex prediction into a sequence of predictions where earlier predictions are used as input to later predictions—commonly referred to as apipeline model. This work describes methods for extending existing margin-based active learning techniques to these two settings, thus increasing the scope of problems for which active learning can be applied. We empirically validate these proposed active learning techniques by reducing the annotated data requirements on multiple instances of synthetic data, a semantic role labeling task, and a named entity and relation extraction system.