Efficient top-n recommendation for very large scale binary rated datasets

Efficient top-n recommendation for very large scale binary rated datasets
复制标题

DOI:
10.1145/2507157.2507189
复制
发表时间:
2013-10
期刊:
Proceedings of the 7th ACM conference on Recommender systems
影响因子:
--
通讯作者:
F. Aiolli
F. Aiolli
中科院分区:
其他
文献类型:
--
作者:
F. Aiolli

文献摘要

被引文献

相似文献

我们提出了一种简单且可扩展的top-N推荐算法,能够处理非常大的数据集和(二进制评级)隐式反馈。我们专注于基于记忆的协同过滤算法,类似于众所周知的基于邻居的显式反馈技术。使该算法具有特别可扩展性的主要区别在于,它只使用正反馈,而不需要执行完整的(用户对用户或项对项)相似性矩阵的显式计算。该算法的研究是在百万歌曲数据集(MSD)挑战的数据上进行的,该挑战的任务是在给定一半的用户收听历史和其他100万人的完整收听历史的情况下,向超过10万用户推荐一组歌曲(从超过38万首可用歌曲中)。特别地,我们研究了整个推荐管道,从定义合适的相似度和评分函数以及如何聚合多个排名策略来定义整体推荐的建议开始。我们提出的技术扩展并改进了去年已经赢得MSD挑战的技术。
We present a simple and scalable algorithm for top-N recommendation able to deal with very large datasets and (binary rated) implicit feedback. We focus on memory-based collaborative filtering algorithms similar to the well known neighboor based technique for explicit feedback. The major difference, that makes the algorithm particularly scalable, is that it uses positive feedback only and no explicit computation of the complete (user-by-user or item-by-item) similarity matrix needs to be performed. The study of the proposed algorithm has been conducted on data from the Million Songs Dataset (MSD) challenge whose task was to suggest a set of songs (out of more than 380k available songs) to more than 100k users given half of the user listening history and complete listening history of other 1 million people. In particular, we investigate on the entire recommendation pipeline, starting from the definition of suitable similarity and scoring functions and suggestions on how to aggregate multiple ranking strategies to define the overall recommendation. The technique we are proposing extends and improves the one that already won the MSD challenge last year.