What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?

What Ingredients Make for an Effective Crowdsourcing Protocol for Difficult NLU Data Collection Tasks?
复制标题

DOI:
10.18653/v1/2021.acl-long.98
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Nikita Nangia;Saku Sugawara;H. Trivedi;Alex Warstadt;Clara Vania;Sam Bowman
Nikita Nangia;Saku Sugawara;H. Trivedi;Alex Warstadt;Clara Vania;Sam Bowman
中科院分区:
其他
文献类型:
--
作者:
Nikita Nangia;Saku Sugawara;H. Trivedi;Alex Warstadt;Clara Vania;Sam Bowman

文献摘要

被引文献

相似文献

众包广泛用于为常见的自然语言理解任务创建数据。尽管这些数据集对于测量和完善语言模型理解很重要,但很少有人关注用于收集数据集的众包方法。在本文中,我们比较了先前工作中提出的作为提高数据质量的方法的干预措施的有效性。我们使用多项选择题回答作为测试平台,并通过指派众包工作者根据四种不同的数据收集协议之一编写问题来进行随机试验。我们发现,要求工作人员为他们的示例编写解释对于提高 NLU 示例难度来说是一种无效的独立策略。然而,我们发现,培训众包工作人员,然后使用收集数据、发送反馈并根据专家判断对工作人员进行资格鉴定的迭代过程,是收集具有挑战性的数据的有效手段。但事实证明,使用众包而不是专家判断来鉴定工人资格并发送反馈并不有效。我们观察到,通过多种措施,来自专家评估的迭代协议的数据更具挑战性。值得注意的是,该数据一致同意部分的人类模型差距平均是基线协议数据差距的两倍。
Crowdsourcing is widely used to create data for common natural language understanding tasks. Despite the importance of these datasets for measuring and refining model understanding of language, there has been little focus on the crowdsourcing methods used for collecting the datasets. In this paper, we compare the efficacy of interventions that have been proposed in prior work as ways of improving data quality. We use multiple-choice question answering as a testbed and run a randomized trial by assigning crowdworkers to write questions under one of four different data collection protocols. We find that asking workers to write explanations for their examples is an ineffective stand-alone strategy for boosting NLU example difficulty. However, we find that training crowdworkers, and then using an iterative process of collecting data, sending feedback, and qualifying workers based on expert judgments is an effective means of collecting challenging data. But using crowdsourced, instead of expert judgments, to qualify workers and send feedback does not prove to be effective. We observe that the data from the iterative protocol with expert assessments is more challenging by several measures. Notably, the human–model gap on the unanimous agreement portion of this data is, on average, twice as large as the gap for the baseline protocol data.