Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations

Interface Design for Crowdsourcing Hierarchical Multi-Label Text Annotations
复制标题

DOI:
10.1145/3544548.3581431
复制
发表时间:
2023-02
期刊:
Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems
影响因子:
--
通讯作者:
Rickard Stureborg;Bhuwan Dhingra;Jun Yang
Rickard Stureborg;Bhuwan Dhingra;Jun Yang
中科院分区:
其他
文献类型:
--
作者:
Rickard Stureborg;Bhuwan Dhingra;Jun Yang

文献摘要

相似文献

人类数据标记是监督学习系统核心的一项重要而昂贵的任务。层次结构帮助人类理解和组织概念。我们问概念层次结构是否以及如何通知注释接口的设计,以提高标签的质量和效率。我们通过对疫苗错误信息的注释来研究这个问题,其中标记任务是困难的并且高度主观的。我们通过收集超过18,000个单独的注释来调查6个用于众包分层标签的用户界面设计。在固定的预算下,将层次结构集成到设计中可以提高众包工作者的F1分数。我们将其归因于:(1)相似的概念,使F1分数比随机分组提高了+0.16;(2)在高难度示例上的相对表现较强(相对F1分数差异为+0.40);(3)过滤掉明显的负面影响,使精度提高了+0.07。最终,整合层次的标签方案优于那些没有实现平均F1为0.70的方案。
Human data labeling is an important and expensive task at the heart of supervised learning systems. Hierarchies help humans understand and organize concepts. We ask whether and how concept hierarchies can inform the design of annotation interfaces to improve labeling quality and efficiency. We study this question through annotation of vaccine misinformation, where the labeling task is difficult and highly subjective. We investigate 6 user interface designs for crowdsourcing hierarchical labels by collecting over 18,000 individual annotations. Under a fixed budget, integrating hierarchies into the design improves crowdsource workers’ F1 scores. We attribute this to (1) Grouping similar concepts, improving F1 scores by +0.16 over random groupings, (2) Strong relative performance on high-difficulty examples (relative F1 score difference of +0.40), and (3) Filtering out obvious negatives, increasing precision by +0.07. Ultimately, labeling schemes integrating the hierarchy outperform those that do not — achieving mean F1 of 0.70.