Unbiased offline recommender evaluation for missing-not-at-random implicit feedback

Unbiased offline recommender evaluation for missing-not-at-random implicit feedback
复制标题

DOI:
10.1145/3240323.3240355
复制
发表时间:
2018-09
期刊:
Proceedings of the 12th ACM Conference on Recommender Systems
影响因子:
--
通讯作者:
Longqi Yang;Yin Cui;Yuan Xuan;Chenyang Wang;Serge J. Belongie;D. Estrin
Longqi Yang;Yin Cui;Yuan Xuan;Chenyang Wang;Serge J. Belongie;D. Estrin
中科院分区:
其他
文献类型:
--
作者:
Longqi Yang;Yin Cui;Yuan Xuan;Chenyang Wang;Serge J. Belongie;D. Estrin

文献摘要

被引文献

相似文献

隐式反馈推荐器(ImplicitRec)只利用积极的用户-项目交互(如点击)来学习个性化的用户偏好。推荐者通常使用从在线平台收集的数据集进行离线评估和比较。这些平台受到流行偏见的影响(即,受欢迎的项目更有可能被呈现并与之交互),因此记录的地面实况数据是非随机缺失(MNAR)。因此,广泛使用的平均总体(AOA)评估器偏向于准确地推荐时尚商品。在本文中,我们(a)AOA的评估偏差和(B)开发一个无偏的和实用的离线评估隐式MNAR数据集使用逆倾向评分(IPS)技术。通过广泛的实验,使用四个真实世界的数据集和四个广泛使用的算法,我们表明:(a)流行的偏见是广泛表现在项目的介绍和互动;(B)由于MNAR数据的评价偏差普遍存在于大多数情况下,AOA是用来评估ImplicitRec;和(c)无偏估计显着减少AOA评价偏差超过30%,在雅虎!平均绝对误差(MAE)。
Implicit-feedback Recommenders (ImplicitRec) leverage positive only user-item interactions, such as clicks, to learn personalized user preferences. Recommenders are often evaluated and compared offline using datasets collected from online platforms. These platforms are subject to popularity bias (i.e., popular items are more likely to be presented and interacted with), and therefore logged ground truth data are Missing-Not-At-Random (MNAR). As a result, the widely used Average-Over-All (AOA) evaluator is biased toward accurately recommending trendy items. In this paper, we (a) investigate evaluation bias of AOA and (b) develop an unbiased and practical offline evaluator for implicit MNAR datasets using the Inverse-Propensity-Scoring (IPS) technique. Through extensive experiments using four real-world datasets and four widely used algorithms, we show that (a) popularity bias is widely manifested in item presentation and interaction; (b) evaluation bias due to MNAR data pervasively exists in most cases where AOA is used to evaluate ImplicitRec; and (c) the unbiased estimator significantly reduces the AOA evaluation bias by more than 30% in the Yahoo! music dataset in terms of the Mean Absolute Error (MAE).