Research-paper recommender systems: a literature survey

Research-paper recommender systems: a literature survey
复制标题

DOI:
10.1007/s00799-015-0156-0
复制
发表时间:
2016-11-01
影响因子:
1.5
通讯作者:
Breitinger, Corinna
Breitinger, Corinna
中科院分区:
其他
文献类型:
--
作者:
Beel, Joeran;Gipp, Bela;Breitinger, Corinna

文献摘要

被引文献

相似文献

在过去的16年里,发表了200多篇关于研究论文推荐系统的研究论文。我们回顾了这些文章,并在本文中提供了一些描述性统计数据,讨论了主要的进步和缺点,并概述了最常见的推荐概念和方法。我们发现超过一半的推荐方法应用了基于内容的过滤(55%)。在被审查的方法中,只有18%的方法采用了协同过滤,16%的方法采用了基于图表的推荐。其他推荐概念包括刻板印象、以项目为中心的推荐和混合推荐。基于内容的过滤方法主要利用用户撰写、标记、浏览或下载的论文。TF-IDF是最常用的加权方案。除了简单的术语外,还使用n-gram、主题和引用来对用户的信息需求建模。我们的综述揭示了当前研究的一些不足之处。首先,目前还不清楚哪些推荐概念和方法最有前途。例如,研究人员报告了基于内容的过滤和协同过滤性能的不同结果。有时基于内容的过滤比协同过滤表现得更好,有时表现得更差。我们确定了结果不明确的三个潜在原因。(A)若干评价有局限性。它们基于经过严格修剪的数据集,很少有用户研究参与者,或者没有使用适当的基线。(B)一些作者提供的关于他们算法的信息很少,这使得重新实现这些方法很困难。因此,研究人员使用相同的建议方法的不同实现,这可能导致结果的变化。(C)我们推测数据集、算法或用户群体的微小变化不可避免地会导致方法性能的强烈变化。因此,找到最有希望的方法是一项挑战。作为第二个限制,我们注意到许多作者忽略了考虑准确性以外的因素,例如总体用户满意度。此外,大多数方法(81%)忽略了用户建模过程,没有自动推断信息,而是让用户提供关键字、文本片段或单个论文作为输入。10%的方法提供了运行时信息。最后,在实践中,很少有研究论文对研究论文推荐系统产生影响。我们还发现该领域缺乏权威和长期研究兴趣:73%的作者在研究论文推荐系统上发表的论文不超过一篇,不同的共同作者群体之间几乎没有合作。我们得出结论,有几项行动可以改善研究前景:开发一个共同的评估框架,就研究论文中包含的信息达成一致,更加关注非准确性方面和用户建模,为研究人员提供一个交换信息的平台,以及一个捆绑可用推荐方法的开源框架。
In the last 16 years, more than 200 research articles were published about research-paper recommender systems. We reviewed these articles and present some descriptive statistics in this paper, as well as a discussion about the major advancements and shortcomings and an overview of the most common recommendation concepts and approaches. We found that more than half of the recommendation approaches applied content-based filtering (55 %). Collaborative filtering was applied by only 18% of the reviewed approaches, and graph-based recommendations by 16%. Other recommendation concepts included stereotyping, item-centric recommendations, and hybrid recommendations. The content-based filtering approaches mainly utilized papers that the users had authored, tagged, browsed, or downloaded. TF-IDF was the most frequently applied weighting scheme. In addition to simple terms, n-grams, topics, and citations were utilized to model users' information needs. Our review revealed some shortcomings of the current research. First, it remains unclear which recommendation concepts and approaches are the most promising. For instance, researchers reported different results on the performance of content-based and collaborative filtering. Sometimes content-based filtering performed better than collaborative filtering and sometimes it performed worse. We identified three potential reasons for the ambiguity of the results. (A) Several evaluations had limitations. They were based on strongly pruned datasets, few participants in user studies, or did not use appropriate baselines. (B) Some authors provided little information about their algorithms, which makes it difficult to re-implement the approaches. Consequently, researchers use different implementations of the same recommendations approaches, which might lead to variations in the results. (C) We speculated that minor variations in datasets, algorithms, or user populations inevitably lead to strong variations in the performance of the approaches. Hence, finding the most promising approaches is a challenge. As a second limitation, we noted that many authors neglected to take into account factors other than accuracy, for example overall user satisfaction. In addition, most approaches (81%) neglected the user-modeling process and did not infer information automatically but let users provide keywords, text snippets, or a single paper as input. Information on runtime was provided for 10% of the approaches. Finally, few research papers had an impact on research-paper recommender systems in practice. We also identified a lack of authority and long-term research interest in the field: 73% of the authors published no more than one paper on research-paper recommender systems, and there was little cooperation among different co-author groups. We concluded that several actions could improve the research landscape: developing a common evaluation framework, agreement on the information to include in research papers, a stronger focus on non-accuracy aspects and user modeling, a platform for researchers to exchange information, and an open-source framework that bundles the available recommendation approaches.