Revisiting Negative Sampling vs. Non-sampling in Implicit Recommendation

Revisiting Negative Sampling vs. Non-sampling in Implicit Recommendation
复制标题

DOI:
10.1145/3522672
复制
发表时间:
2022-03
影响因子:
5.6
通讯作者:
C. Chen;Weizhi Ma;M. Zhang;Chenyang Wang;Yiqun Liu;Shaoping Ma
C. Chen;Weizhi Ma;M. Zhang;Chenyang Wang;Yiqun Liu;Shaoping Ma
中科院分区:
计算机科学2区
文献类型:
--
作者:
C. Chen;Weizhi Ma;M. Zhang;Chenyang Wang;Yiqun Liu;Shaoping Ma

文献摘要

相似文献

推荐系统在缓解信息过载问题方面发挥着重要作用。通常,推荐模型被训练为辨别每个用户的正面(喜欢)和负面(不喜欢)实例。然而,在开放世界假设下,只有积极的情况下,没有负面的情况下,用户的隐式反馈,这造成了缺乏负样本的不平衡学习的挑战。为了解决这个问题,之前已经提出了两种类型的学习策略,负采样策略和非采样策略。第一种策略从缺失数据中采样负实例(即,未标记数据),而非抽样策略将所有缺失数据视为阴性。虽然学习策略是算法性能的关键,但负采样和非采样的深入比较还没有得到充分的研究。为了弥补这一差距,我们系统地分析了负采样和非采样的隐式推荐在这项工作中的作用。具体而言,我们首先从理论上重新审视了负抽样和不抽样的异议。然后,通过仔细设置各种代表性的推荐方法,我们探讨了负采样和非采样在不同场景下的性能。我们的研究结果实证表明,虽然负采样已被广泛应用于最近的推荐模型,它是不平凡的均匀采样方法,以显示可比的性能非采样学习方法。最后,我们讨论了负采样和非采样的可扩展性和复杂性,并提出了一些开放的问题和未来的研究课题,值得进一步探讨。
Recommendation systems play an important role in alleviating the information overload issue. Generally, a recommendation model is trained to discern between positive (liked) and negative (disliked) instances for each user. However, under the open-world assumption, there are only positive instances but no negative instances from users’ implicit feedback, which poses the imbalanced learning challenge of lacking negative samples. To address this, two types of learning strategies have been proposed before, the negative sampling strategy and non-sampling strategy. The first strategy samples negative instances from missing data (i.e., unlabeled data), while the non-sampling strategy regards all the missing data as negative. Although learning strategies are known to be essential for algorithm performance, the in-depth comparison of negative sampling and non-sampling has not been sufficiently explored by far. To bridge this gap, we systematically analyze the role of negative sampling and non-sampling for implicit recommendation in this work. Specifically, we first theoretically revisit the objection of negative sampling and non-sampling. Then, with a careful setup of various representative recommendation methods, we explore the performance of negative sampling and non-sampling in different scenarios. Our results empirically show that although negative sampling has been widely applied to recent recommendation models, it is non-trivial for uniform sampling methods to show comparable performance to non-sampling learning methods. Finally, we discuss the scalability and complexity of negative sampling and non-sampling and present some open problems and future research topics that are worth being further explored.