Predicting Long-Term Citations from Short-Term Linguistic Influence

Predicting Long-Term Citations from Short-Term Linguistic Influence
复制标题

DOI:
10.48550/arxiv.2210.13628
复制
发表时间:
2022-10
期刊:
--
影响因子:
--
通讯作者:
Sandeep Soni;David Bamman;Jacob Eisenstein
Sandeep Soni;David Bamman;Jacob Eisenstein
中科院分区:
其他
文献类型:
--
作者:
Sandeep Soni;David Bamman;Jacob Eisenstein

文献摘要

相似文献

衡量一篇研究论文影响力的标准是它被引用的次数。然而,论文被引用的原因有很多,引用计数提供的关于一篇论文对后续出版物内容影响程度的信息有限。因此,我们提出了一种新的方法来量化语言的影响,在时间戳的文档集合。有两个主要步骤:首先,使用上下文嵌入和词频识别词汇和语义变化;其次,通过估计具有低秩参数矩阵的高维Hawkes过程,将关于这些变化的信息聚合到每个文档的影响分数中。我们发现,这种语言影响力的措施是预测$\textit{future}$引用:语言影响力的估计,从两年后的论文的出版是相关的,并预测其引用计数在接下来的三年。这是证明使用增量时间训练/测试分裂的在线评估,与一个强大的基线,包括初始引用计数,主题和词汇功能的预测。
A standard measure of the influence of a research paper is the number of times it is cited. However, papers may be cited for many reasons, and citation count offers limited information about the extent to which a paper affected the content of subsequent publications. We therefore propose a novel method to quantify linguistic influence in timestamped document collections. There are two main steps: first, identify lexical and semantic changes using contextual embeddings and word frequencies; second, aggregate information about these changes into per-document influence scores by estimating a high-dimensional Hawkes process with a low-rank parameter matrix. We show that this measure of linguistic influence is predictive of $\textit{future}$ citations: the estimate of linguistic influence from the two years after a paper's publication is correlated with and predictive of its citation count in the following three years. This is demonstrated using an online evaluation with incremental temporal training/test splits, in comparison with a strong baseline that includes predictors for initial citation counts, topics, and lexical features.