Optimal Budget Allocation for Crowdsourcing Labels for Graphs

Optimal Budget Allocation for Crowdsourcing Labels for Graphs
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Adithya Kulkarni;Mohna Chakraborty;Sihong Xie;Qi Li
Adithya Kulkarni;Mohna Chakraborty;Sihong Xie;Qi Li
中科院分区:
其他
文献类型:
--
作者:
Adithya Kulkarni;Mohna Chakraborty;Sihong Xie;Qi Li

文献摘要

相似文献

众包是一种有效和高效的范式,用于为雇用人群工作者的未标记语料库获取标签。这项工作认为,预算分配问题的一个广义的设置上的图形的实例被标记的边缘编码实例依赖关系。具体来说,给定一个图和一个标记预算,我们提出了一个最优策略来在实例之间分配预算,以最大化整体标记准确度。我们将问题表述为贝叶斯马尔可夫决策过程(MDP),其中我们将任务定义为在预算约束下最大化整体标签准确性的优化问题。然后,我们提出了一种新的阶段式奖励函数,该函数考虑了每个时间戳工人标签对整个图的影响。该奖励函数用于为优化问题找到最优策略。从理论上讲,我们表明,当预算有限时,我们提出的政策是一致的。我们在五个真实世界的图数据集上进行了广泛的实验,并证明了所提出的策略在预算限制下实现更高标签准确性的有效性。
Crowdsourcing is an effective and efficient paradigm for obtaining labels for unlabeled corpus employing crowd workers. This work considers the budget allocation problem for a generalized setting on a graph of instances to be labeled where edges encode instance dependencies. Specifically, given a graph and a labeling budget, we propose an optimal policy to allocate the budget among the instances to maximize the overall labeling accuracy. We formulate the problem as a Bayesian Markov Decision Process (MDP), where we define our task as an optimization problem that maximizes the overall label accuracy under budget constraints. Then, we propose a novel stage-wise reward function that considers the effect of worker labels on the whole graph at each timestamp. This reward function is utilized to find an optimal policy for the optimization problem. Theoretically, we show that our proposed policies are consistent when the budget is infinite. We conduct extensive experiments on five real-world graph datasets and demonstrate the effectiveness of the proposed policies to achieve a higher label accuracy under budget constraints.