Crowdsourcing Human Annotation on Web Page Structure: Infrastructure Design and Behavior-Based Quality Control

Crowdsourcing Human Annotation on Web Page Structure: Infrastructure Design and Behavior-Based Quality Control
复制标题

DOI:
10.1145/2870649
复制
发表时间:
2016-07-01
影响因子:
5
通讯作者:
Huynh, David
Huynh, David
中科院分区:
计算机科学3区
文献类型:
--
作者:
Han, Shuguang;Dai, Peng;Huynh, David

文献摘要

被引文献

相似文献

分析网页的语义结构是网页信息抽取的关键组成部分。成功的抽取算法通常需要大规模的训练和评估数据集,而这些数据集很难获得。最近,众包已被证明是在不需要太多领域知识的领域收集大规模训练数据的一种有效方法。对于更复杂的领域,研究人员提出了复杂的质量控制机制,以并行或顺序的方式复制任务,然后聚合来自多个工作者的响应。传统的标注集成方法往往更信任历史性能较高的工作者,因此被称为基于性能的方法。最近,Rzeszotarski和Kitture已经证明,在几个众包应用中,行为特征也与注释质量高度相关。在这篇文章中,我们提出了一个新的众包系统,称为Wernicke,为网络信息提取提供注释。Wernicke收集了广泛的行为特征,并基于这些特征预测了一个具有挑战性的任务领域的注释质量:注释网页结构。我们通过一个案例研究来评估使用行为特征进行质量控制的有效性,其中32名工作人员标注了来自五个流行网站的200个问答网页。在这样做的过程中,我们发现了几点:(1)许多行为特征对众包质量有显著的预测作用。(2)基于行为特征的方法在召回预测方面优于基于性能的方法,而在查准率预测方面与基于行为特征的方法相当。此外,使用行为特征不太容易受到冷启动问题的影响,并且相应的预测模型对于预测召回比用于跨网站质量分析的精度更具普适性。(3)可以有效地将员工的行为信息和历史绩效信息结合起来,进一步减少预测误差。
Parsing the semantic structure of a web page is a key component of web information extraction. Successful extraction algorithms usually require large-scale training and evaluation datasets, which are difficult to acquire. Recently, crowdsourcing has proven to be an effective method of collecting large-scale training data in domains that do not require much domain knowledge. For more complex domains, researchers have proposed sophisticated quality control mechanisms to replicate tasks in parallel or sequential ways and then aggregate responses from multiple workers. Conventional annotation integration methods often put more trust in the workers with high historical performance; thus, they are called performance-based methods. Recently, Rzeszotarski and Kittur have demonstrated that behavioral features are also highly correlated with annotation quality in several crowdsourcing applications. In this article, we present a new crowdsourcing system, called Wernicke, to provide annotations for web information extraction. Wernicke collects a wide set of behavioral features and, based on these features, predicts annotation quality for a challenging task domain: annotating web page structure. We evaluate the effectiveness of quality control using behavioral features through a case study where 32 workers annotate 200 Q&A web pages from five popular websites. In doing so, we discover several things: (1) Many behavioral features are significant predictors for crowdsourcing quality. (2) The behavioral-feature-based method outperforms performance-based methods in recall prediction, while performing equally with precision prediction. In addition, using behavioral features is less vulnerable to the cold-start problem, and the corresponding prediction model is more generalizable for predicting recall than precision for cross-website quality analysis. (3) One can effectively combine workers' behavioral information and historical performance information to further reduce prediction errors.