Learning a Cost-Effective Annotation Policy for Question Answering

Learning a Cost-Effective Annotation Policy for Question Answering
复制标题

DOI:
10.18653/v1/2020.emnlp-main.246
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Bernhard Kratzwald;S. Feuerriegel;Huan Sun
Bernhard Kratzwald;S. Feuerriegel;Huan Sun
中科院分区:
其他
文献类型:
--
作者:
Bernhard Kratzwald;S. Feuerriegel;Huan Sun

文献摘要

相似文献

最先进的问题回答(QA)依赖于大量的训练数据,标记是耗时的,因此是昂贵的。因此,定制QA系统具有挑战性。作为一种补救措施,我们提出了一种新的框架来注释QA数据集,需要学习一个具有成本效益的注释策略和半监督注释方案。后者减少了人工工作:它利用底层QA系统来建议潜在的候选注释。然后,人类注释器简单地提供关于这些候选项的二进制反馈。我们的系统被设计为使过去的注释不断提高未来的性能,从而提高整体注释成本。据我们所知,这是第一篇以最小的注释成本来解决注释问题的论文。我们比较我们的框架对传统的手动注释在一组广泛的实验。我们发现我们的方法可以减少高达21.1%的注释成本。
State-of-the-art question answering (QA) relies upon large amounts of training data for which labeling is time consuming and thus expensive. For this reason, customizing QA systems is challenging. As a remedy, we propose a novel framework for annotating QA datasets that entails learning a cost-effective annotation policy and a semi-supervised annotation scheme. The latter reduces the human effort: it leverages the underlying QA system to suggest potential candidate annotations. Human annotators then simply provide binary feedback on these candidates. Our system is designed such that past annotations continuously improve the future performance and thus overall annotation cost. To the best of our knowledge, this is the first paper to address the problem of annotating questions with minimal annotation cost. We compare our framework against traditional manual annotations in an extensive set of experiments. We find that our approach can reduce up to 21.1% of the annotation cost.