An analysis of the 2014 RecSys Challenge

An analysis of the 2014 RecSys Challenge
复制标题

DOI:
10.1145/2668067.2668082
复制
发表时间:
2014-10
期刊:
--
影响因子:
--
通讯作者:
D. Loiacono;A. Lommatzsch;R. Turrin
D. Loiacono;A. Lommatzsch;R. Turrin
中科院分区:
其他
文献类型:
--
作者:
D. Loiacono;A. Lommatzsch;R. Turrin

文献摘要

相似文献

2014年RecSys挑战赛的重点是智能手机IMDb应用程序用户发布的推文所产生的参与度。这种参与取决于以下属性:发布消息的用户(例如,他在社交网络中的角色)、推文内容(例如,评级)、以及推文的电影对象(例如,电影的受欢迎程度)。在这项工作中,我们提供了对数据集和任务的分析,以帮助参与者更好地理解挑战。此外,我们还提出了一种基线预测算法。我们将我们的分析分为三个阶段:(I)数据丰富,(Ii)知识提取,和(Iii)参与度预测。最初,我们使用从IMDb和Free-base中提取的其他电影属性来丰富数据集。随后,我们分析了主要推文属性的统计数据,并应用了一些机器学习技术来从数据中提取额外的知识。最后,我们根据这些分析的主要结果定义了一个预测因子。我们定义了一个关于属性的线性回归模型,例如:用户评分、提及的存在,以及这条推文是转发还是已经被转发。这样的预测导致nDCG@10等于0.8352。
The RecSys challenge 2014 focuses on the engagement generated by the tweets posted by the users of the IMDb application for smartphones. Such engagement depends on attributes concerning: the user who posts the message (e.g., his role in the social network), the tweet content (e.g., the rating), and the movie object of the tweet (e.g., the popularity of the movie). In this work we provide an analysis of the dataset and of the task to help participants better understand the challenge. Furthermore, we propose a baseline prediction algorithm. We split our analysis into three stages: (i) data enrichment, (ii) knowledge extraction, and (iii) engagement prediction. Initially, we enriched the dataset with additional movie attributes extracted from IMDb and Free-base. Subsequently, we analyzed the statistics of the main tweet attributes and we applied some machine learning techniques to extract additional knowledge from the data. Finally, we defined a predictor on the basis of the main outcomes of these analyses. We define a linear regression model on attributes such as: user rating score, the presence of mentions, and whether the tweet is a retweet or it has already been retweeted. Such predictor led to an nDCG@10 equals to 0.8352.