Large Scale Subject Category Classification of Scholarly Papers With Deep Attentive Neural Networks.

Large Scale Subject Category Classification of Scholarly Papers With Deep Attentive Neural Networks.
复制标题

DOI:
10.3389/frma.2020.600382
复制
发表时间:
2020
影响因子:
--
通讯作者:
Giles CL
Giles CL
中科院分区:
其他
文献类型:
--
作者:
Kandimalla B;Rohatgi S;Wu J;Giles CL

文献摘要

被引文献

相似文献

学术论文的主题类别通常是指论文所属的知识领域,例如计算机科学或物理学。主题类别分类是文献计量学研究、组织科学出版物以提取领域知识和促进数字图书馆搜索引擎的分面搜索的先决条件。不幸的是,许多学术论文没有将这些信息作为其元数据的一部分。解决这一任务的大多数现有方法都集中在无监督学习上,通常依赖于引文网络。然而,引用当前论文的完整论文列表可能不容易获得。特别是,很少引用或没有引用的新论文无法使用这种方法进行分类。在这里,我们提出了一个深度关注神经网络(DANN),该网络仅使用摘要对学术论文进行分类。该网络使用来自Web of Science (WoS)的900万篇摘要进行训练。我们还使用涵盖104个主题类别的WoS模式。该网络由两个双向递归神经网络和一个注意层组成。我们通过改变架构和文本表示将模型与基线进行比较。我们的最佳模型达到了0.76的微观测量,单个主题类别的范围在0.50到0.95之间。结果表明,再训练词嵌入模型对于最大化词汇重叠的重要性和注意机制的有效性。单词向量与TFIDF的结合优于字符级和句子级嵌入模型。我们讨论了不平衡的样本和重叠的类别,并提出了可能的缓解策略。我们还通过对100万篇学术论文的随机样本进行分类来确定CiteSeerX中的主题类别分布。
Subject categories of scholarly papers generally refer to the knowledge domain(s) to which the papers belong, examples being computer science or physics. Subject category classification is a prerequisite for bibliometric studies, organizing scientific publications for domain knowledge extraction, and facilitating faceted searches for digital library search engines. Unfortunately, many academic papers do not have such information as part of their metadata. Most existing methods for solving this task focus on unsupervised learning that often relies on citation networks. However, a complete list of papers citing the current paper may not be readily available. In particular, new papers that have few or no citations cannot be classified using such methods. Here, we propose a deep attentive neural network (DANN) that classifies scholarly papers using only their abstracts. The network is trained using nine million abstracts from Web of Science (WoS). We also use the WoS schema that covers 104 subject categories. The proposed network consists of two bi-directional recurrent neural networks followed by an attention layer. We compare our model against baselines by varying the architecture and text representation. Our best model achieves micro- measure of 0.76 with of individual subject categories ranging from 0.50 to 0.95. The results showed the importance of retraining word embedding models to maximize the vocabulary overlap and the effectiveness of the attention mechanism. The combination of word vectors with TFIDF outperforms character and sentence level embedding models. We discuss imbalanced samples and overlapping categories and suggest possible strategies for mitigation. We also determine the subject category distribution in CiteSeerX by classifying a random sample of one million academic papers.