Meta-Path-Based Ranking with Pseudo Relevance Feedback on Heterogeneous Graph for Citation Recommendation

Meta-Path-Based Ranking with Pseudo Relevance Feedback on Heterogeneous Graph for Citation Recommendation
复制标题

DOI:
10.1145/2661829.2661965
复制
发表时间:
2014-11
期刊:
Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Xiaozhong Liu;Yingying Yu;Chun Guo;Yizhou Sun
Xiaozhong Liu;Yingying Yu;Chun Guo;Yizhou Sun
中科院分区:
其他
文献类型:
--
作者:
Xiaozhong Liu;Yingying Yu;Chun Guo;Yizhou Sun

文献摘要

被引文献

相似文献

在线提供的学术出版物数量庞大,这对学者如何检索可用的新信息和找到候选参考论文提出了重大挑战。虽然经典的文本检索和伪相关反馈(PRF)算法可以帮助学者访问所需的出版物,但在本研究中,我们通过利用异构书目图上的许多元路径,提出了一种基于PRF的创新出版物排名方法。图上的不同元路径解决不同的排名假设,而伪相关论文(来自检索结果)被用作图上的种子节点。同时,与以往的研究不同,我们提出了一个新的上下文丰富的异构网络提取全文出版物内容沿着与引用上下文促进“限制元路径”。通过使用learning-to-rank,我们整合了18种不同的基于元路径的排名特征,以获得候选被引论文的最终排名分数。在ACM全文语料库上的实验结果表明,在新的图上,基于元路径的PRF排序算法显著优于基于文本或基于PageRank的PRF的文本检索算法(p < 0.0001)。
The sheer volume of scholarly publications available online significantly challenges how scholars retrieve the new information available and locate the candidate reference papers. While classical text retrieval and pseudo relevance feedback (PRF) algorithms can assist scholars in accessing needed publications, in this study, we propose an innovative publication ranking method with PRF by leveraging a number of meta-paths on the heterogeneous bibliographic graph. Different meta-paths on the graph address different ranking hypotheses, whereas the pseudo-relevant papers (from the retrieval results) are used as the seed nodes on the graph. Meanwhile, unlike prior studies, we propose "restricted meta-path" facilitated by a new context-rich heterogeneous network extracted from full-text publication content along with citation context. By using learning-to-rank, we integrate 18 different meta-path-based ranking features to derive the final ranking scores for candidate cited papers. Experimental results with ACM full-text corpus show that meta-path-based ranking with PRF on the new graph significantly (p < 0.0001) outperforms text retrieval algorithms with text-based or PageRank-based PRF.