HINTS: Citation Time Series Prediction for New Publications via Dynamic Heterogeneous Information Network Embedding

HINTS: Citation Time Series Prediction for New Publications via Dynamic Heterogeneous Information Network Embedding
复制标题

DOI:
10.1145/3442381.3450107
复制
发表时间:
2021-04
期刊:
Proceedings of the Web Conference 2021
影响因子:
--
通讯作者:
Song Jiang;Bernard Koch;Yizhou Sun
Song Jiang;Bernard Koch;Yizhou Sun
中科院分区:
其他
文献类型:
--
作者:
Song Jiang;Bernard Koch;Yizhou Sun

文献摘要

相似文献

科学影响力的准确预测对于科学家、学术推荐系统和资助机构都很重要。现有的方法依赖于多年的领先引用值来预测科学论文的引用(影响力的代理),即使大多数论文在发表后的头几年中做出了最大的贡献。在本文中,我们解决了一个新的问题:预测一篇新论文从发表之日起的引用时间序列(即,没有前导值)。我们提出了HINTS,这是一种新型的端到端深度学习框架,可以将动态异构信息网络(DHIN)中的引用信号转换为引用时间序列。HINTS在论文发表前几年从DHIN嵌入中估算出伪领先值,然后将这些嵌入转换为正式模型的参数,该模型可以在发表后立即预测引用计数。对计算机科学和物理学两个真实数据集的实证分析表明,HINTS与基线引文预测模型具有竞争力。虽然我们专注于引用,但我们的方法可以推广到其他“冷启动”时间序列预测任务,其中关系数据可用,并且早期时间戳的准确预测至关重要。
Accurate prediction of scientific impact is important for scientists, academic recommender systems, and granting organizations alike. Existing approaches rely on many years of leading citation values to predict a scientific paper’s citations (a proxy for impact), even though most papers make their largest contributions in the first few years after they are published. In this paper, we tackle a new problem: predicting a new paper’s citation time series from the date of publication (i.e., without leading values). We propose HINTS, a novel end-to-end deep learning framework that converts citation signals from dynamic heterogeneous information networks (DHIN) into citation time series. HINTS imputes pseudo-leading values for a paper in the years before it is published from DHIN embeddings, and then transforms these embeddings into the parameters of a formal model that can predict citation counts immediately after publication. Empirical analysis on two real-world datasets from Computer Science and Physics show that HINTS is competitive with baseline citation prediction models. While we focus on citations, our approach generalizes to other “cold start” time series prediction tasks where relational data is available and accurate prediction in early timestamps is crucial.