A Bayesian Approach for Sequence Tagging with Crowds

A Bayesian Approach for Sequence Tagging with Crowds
复制标题

群体序列标记的贝叶斯方法

DOI:
10.18653/v1/d19-1101
复制
发表时间:
2018
期刊:
SIAM J. Comput.
影响因子:
--
通讯作者:
Iryna Gurevych
Iryna Gurevych
中科院分区:
--
文献类型:
--
作者:
Edwin Simpson;Iryna Gurevych

文献摘要

参考文献

被引文献

相似文献

目前用于序列标记的方法是NLP中的一项核心任务,需要大量的数据,这促使人们使用众包作为获得标记数据的一种廉价方式。然而,注释器通常是不可靠的,并且当前的聚集方法不能捕获常见类型的SPAN注解错误。为了解决这个问题,我们提出了一种聚集序列标签的贝叶斯方法,该方法通过对注释和地面事实标签之间的顺序依赖进行建模来减少错误。通过采用贝叶斯方法,我们考虑了模型中的不确定性,这是由于注释器错误和缺乏用于建模注释器的数据,这些注释器完成的任务很少。我们在用于命名实体识别、信息提取和参数挖掘的众包数据上对我们的模型进行了评估,结果表明我们的序列模型的性能优于以前的技术水平,贝叶斯方法的性能优于非贝叶斯方法。我们还发现,我们的方法可以通过更有效的主动学习来降低众包成本,因为它在注释较少的情况下更好地捕获了序列标签中的不确定性。
Current methods for sequence tagging, a core task in NLP, are data hungry, which motivates the use of crowdsourcing as a cheap way to obtain labelled data. However, annotators are often unreliable and current aggregation methods cannot capture common types of span annotation error. To address this, we propose a Bayesian method for aggregating sequence tags that reduces errors by modelling sequential dependencies between the annotations as well as the ground-truth labels. By taking a Bayesian approach, we account for uncertainty in the model due to both annotator errors and the lack of data for modelling annotators who complete few tasks. We evaluate our model on crowdsourced data for named entity recognition, information extraction and argument mining, showing that our sequential model outperforms the previous state of the art, and that Bayesian approaches outperform non-Bayesian alternatives. We also find that our approach can reduce crowdsourcing costs through more effective active learning, as it better captures uncertainty in the sequence labels when there are few annotations.
DOI: 10.1007/3-540-45014-9
发表时间: 2000-06
期刊: --
影响因子: --
作者:
Thomas G. Dietterich
通讯作者: Thomas G. Dietterich