Collaborative Filtering with Network Representation Learning for Citation Recommendation

Collaborative Filtering with Network Representation Learning for Citation Recommendation
复制标题

协同过滤与网络表示学习用于引文推荐

DOI:
10.1109/tbdata.2020.3034976
复制
发表时间:
2020
影响因子:
7.2
通讯作者:
Huan Liu
Huan Liu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wei Wang;Tao Tang;Feng Xia;Zhiguo Gong;Zhikui Chen;Huan Liu

文献摘要

相似文献

引文推荐在学术大数据背景下发挥着重要作用,由于信息过载,查找相关论文变得更加困难。传统的协同过滤(CF)应用于引文推荐是具有挑战性的,由于冷启动问题和缺乏论文评级。为了解决这些挑战,在这篇文章中,我们提出了一个协同过滤与网络表示学习框架的引文推荐,即CNCRec,这是一个混合的基于用户的CF同时考虑论文内容和网络拓扑结构。它的目的是在异构的学术信息网络中推荐引文。CNCRec基于属性化引文网络表征学习创建论文评级矩阵,其中属性是从论文文本信息中提取的主题。同时,学习属性协作网络的表示,以改善最近邻居的选择。通过利用网络表示学习的力量,与以前的基于上下文感知的网络模型相比,CNCRec能够充分利用整个引文网络拓扑结构。在DBLP和APS数据集上进行的大量实验表明,该方法在准确率、召回率和MRR(平均倒数秩)方面优于最先进的方法。此外,CNCRec可以更好地解决数据稀疏性问题相比,其他基于CF基线。
Citation recommendation plays an important role in the context of scholarly big data, where finding relevant papers has become more difficult because of information overload. Applying traditional collaborative filtering (CF) to citation recommendation is challenging due to the cold start problem and the lack of paper ratings. To address these challenges, in this article, we propose a collaborative filtering with network representation learning framework for citation recommendation, namely CNCRec, which is a hybrid user-based CF considering both paper content and network topology. It aims at recommending citations in heterogeneous academic information networks. CNCRec creates the paper rating matrix based on attributed citation network representation learning, where the attributes are topics extracted from the paper text information. Meanwhile, the learned representations of attributed collaboration network is utilized to improve the selection of nearest neighbors. By harnessing the power of network representation learning, CNCRec is able to make full use of the whole citation network topology compared with previous context-aware network-based models. Extensive experiments on both DBLP and APS datasets show that the proposed method outperforms state-of-the-art methods in terms of precision, recall, and MRR (Mean Reciprocal Rank). Moreover, CNCRec can better solve the data sparsity problem compared with other CF-based baselines.