Incentivizing High Quality User Contributions: New Arm Generation in Bandit Learning

Incentivizing High Quality User Contributions: New Arm Generation in Bandit Learning
复制标题

激励高质量用户贡献:强盗学习中的新一代 Arm

DOI:
--
复制
发表时间:
2018
期刊:
AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Chien
Chien
中科院分区:
--
文献类型:
--
作者:
Yang Liu;Chien

文献摘要

被引文献

相似文献

我们研究了在用户生成的内容平台中激励高质量贡献的问题,在这些平台中,用户以未知的质量顺序到达。我们感兴趣的是设计一种内容显示策略,该策略决定应该选择哪些内容显示给用户,目标是最大化用户体验(即,这个目标自然导致激励高质量贡献和学习未知内容质量的联合问题。为了解决激励问题,我们考虑了一个模型,其中用户在决定是否贡献时具有策略性,并且受到曝光的激励,即,他们的目标是最大限度地增加他们的贡献被查看的次数。从学习的角度来看,我们将内容质量建模为获得积极反馈的概率(例如,喜欢或赞同)。当然,平台需要解决探索(收集所有内容的反馈)和利用(显示最佳内容)之间的经典权衡。我们将此问题表述为多臂强盗问题,其中臂的数量(即,捐款)随着时间的推移而增加,并取决于到达用户的战略选择。我们首先表明,应用标准的强盗算法激励洪水的低成本的贡献,这反过来又导致线性遗憾。然后,我们提出了兰德_UCB,它增加了一个额外的随机化层上的UCB算法,以解决洪水的贡献的问题。我们发现,兰德_UCB有助于消除低质量的贡献的激励,提供高质量的贡献的激励(由于有限数量的探索低质量的),并实现了次线性的遗憾方面显示当前最好的武器。
We study the problem of incentivizing high quality contributions in user generated content platforms, in which users arrive sequentially with unknown quality. We are interested in designing a content displaying strategy which decides which content should be chosen to show to users, with the goal of maximizing user experience (i.e., the likelihood of users liking the content).This goal naturally leads to a joint problem of incentivizing high quality contributions and learning the unknown content quality. To address the incentive issue, we consider a model in which users are strategic in deciding whether to contribute and are motivated by exposure, i.e., they aim to maximize the number of times their contributions are viewed. For the learning perspective, we model the content quality as the probability of obtaining positive feedback (e.g., like or upvote) from a random user. Naturally, the platform needs to resolve the classical trade-off between exploration (collecting feedback for all content) and exploitation (displaying the best content). We formulate this problem as a multi-arm bandit problem, where the number of arms (i.e., contributions) is increasing over time and depends on the strategic choices of arriving users. We first show that applying standard bandit algorithms incentivizes a flood of low cost contributions, which in turn leads to linear regret. We then propose Rand_UCB which adds an additional layer of randomization on top of the UCB algorithm to address the issue of flooding contributions. We show that Rand_UCB helps eliminate the incentives for low quality contributions, provides incentives for high quality contributions (due to bounded number of explorations for the low quality ones), and achieves sub-linear regrets with respect to displaying the current best arms.