Bayesian Decision Process for Budget-efficient Crowdsourced Clustering

Bayesian Decision Process for Budget-efficient Crowdsourced Clustering
复制标题

DOI:
10.24963/ijcai.2020/283
复制
发表时间:
2020-07
期刊:
--
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

集群的性能取决于两个项目之间适当定义的相似性。当基于人类感知来测量相似性时,人类工作人员通常被用来估计项目之间的相似性得分以支持聚类,从而导致称为众包聚类的过程。假设为每个相似性得分向工作者支付金钱奖励,并且假设对之间的相似性和工作者的可靠性具有很大的多样性,当预算有限时,明智地将项目对分配给不同的工作者以优化聚类结果是至关重要的。我们将此预算分配问题建模为马尔可夫决策过程,其中项目对根据其提供的历史相似性得分动态分配给工人。我们提出了一个乐观的知识梯度政策,在每个阶段的项目的分配是基于最小权重K-割定义的相似图。我们提供模拟研究和真实的数据分析,以证明所提出的方法的性能。
The performance of clustering depends on an appropriately defined similarity between two items. When the similarity is measured based on human perception, human workers are often employed to estimate a similarity score between items in order to support clustering, leading to a procedure called crowdsourced clustering. Assuming a monetary reward is paid to a worker for each similarity score and assuming the similarities between pairs and workers' reliability have a large diversity, when the budget is limited, it is critical to wisely assign pairs of items to different workers to optimize the clustering result. We model this budget allocation problem as a Markov decision process where item pairs are dynamically assigned to workers based on the historical similarity scores they provided. We propose an optimistic knowledge gradient policy where the assignment of items in each stage is based on the minimum-weight K-cut defined on a similarity graph. We provide simulation studies and real data analysis to demonstrate the performance of the proposed method.