CRUX: Adaptive Querying for Efficient Crowdsourced Data Extraction

CRUX: Adaptive Querying for Efficient Crowdsourced Data Extraction
复制标题

CRUX:用于高效众包数据提取的自适应查询

DOI:
10.1145/3357384.3357976
复制
发表时间:
2019
期刊:
Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Parameswaran, Aditya
Parameswaran, Aditya
中科院分区:
--
文献类型:
--
作者:
Rekatsinas, Theodoros;Deshpande, Amol;Parameswaran, Aditya

文献摘要

参考文献

相似文献

众包对于收集有关真实世界实体的信息至关重要。现有的众包数据提取解决方案使用固定的、非自适应的查询策略,其重复地要求工作者提供来自固定域的实体,直到达到期望的覆盖水平。不幸的是,这样的解决方案是非常不切实际的,因为它们产生许多重复的提取。我们设计了一个自适应查询框架,CRUX,最大限度地提高了提取的实体的数量为一个给定的预算。我们表明,预算众包实体提取的问题是NP难的。我们利用两种见解来集中我们的提取工作:利用感兴趣的域的结构,以及使用排除列表来限制重复提取。我们开发了新的统计工具,原因\em额外的查询的新的不同的提取实体的数量在存在的信息很少,并将它们嵌入到自适应算法,最大限度地提高不同的提取实体的预算约束下。我们在合成和真实世界的数据集上评估了我们的技术,证明了在相同预算下,与竞争方法相比,改进高达300%。
Crowdsourcing is essential for collecting information about real-world entities. Existing crowdsourced data extraction solutions use fixed, non-adaptive querying strategies that repeatedly ask workers to provide entities from a fixed domain until a desired level of coverage is reached. Unfortunately, such solutions are highly impractical as they yield many duplicate extractions. We design an adaptive querying framework, CRUX, that maximizes the number of extracted entities for a given budget. We show that the problem of budgeted crowdsourced entity extraction is NP-Hard. We leverage two insights to focus our extraction efforts: \em exploiting the structure of the domain of interest, and \em using exclude lists to limit repeated extractions. We develop new statistical tools to reason about the number of new distinct extracted entities of \em additional queries under the presence of little information, and embed them within adaptive algorithms that maximize the distinct extracted entities under budget constraints. We evaluate our techniques on synthetic and real-world datasets, demonstrating an improvement of up to 300% over competing approaches for the same budget.
AskSheet:使用电子表格进行高效的人工计算决策
DOI: 10.1145/2531602.2531728
发表时间: 2014
期刊: Proceedings of the 17th ACM conference on Computer supported cooperative work & social computing
影响因子: --
作者:
Alexander J. Quinn;B. Bederson
通讯作者: B. Bederson
应用于森林群落的物种丰富度小样本估计
DOI: --
发表时间: 2010
期刊: Biometrics
影响因子: 1.9
作者:
W. Hwang;Tsung‐Jen Shen
通讯作者: Tsung‐Jen Shen
DOI: --
发表时间: 2006-12
期刊: J. Mach. Learn. Res.
影响因子: --
作者:
Eyal Even-Dar;Shie Mannor;Y. Mansour
通讯作者: Eyal Even-Dar;Shie Mannor;Y. Mansour
动态过滤器:群体的自适应查询处理
DOI: --
发表时间: 2017
期刊: Fifth AAAI Conference on Human Computation and Crowdsourcing
影响因子: --
作者:
Lan, Doren;Reed, Katherine;Shin, Austin;Trushkowsky, Beth
通讯作者: Trushkowsky, Beth
DOI: --
发表时间: 2009
期刊: ACM SIGACT-SIGMOD-SIGART Symposium on Principles of Database Systems
影响因子: --
作者:
Nilesh N. Dalvi;Ravi Kumar;B. Pang;R. Ramakrishnan;A. Tomkins;Philip Bohannon;S. Keerthi;S. Merugu
通讯作者: S. Merugu