Merging the citations received by arXiv-deposited e-prints and their corresponding published journal articles: Problems and perspectives

Merging the citations received by arXiv-deposited e-prints and their corresponding published journal articles: Problems and perspectives
复制标题

合并 arXiv 存储的电子印刷品收到的引用及其相应发表的期刊文章:问题和观点

DOI:
10.1016/j.ipm.2020.102267
复制
发表时间:
2020-09-01
影响因子:
8.6
通讯作者:
Zhu, Linna
Zhu, Linna
中科院分区:
计算机科学1区
文献类型:
--
作者:
Gao, Yong;Wu, Qiang;Zhu, Linna

文献摘要

被引文献

相似文献

将arXiv存放的电子印刷品(arXiv版本)的引文计数与其相应的已发表期刊文章(出版商版本)的引文计数合并是引文分析中的一个重要问题。本文以arXiv存储的电子印刷品为例,采用人工方法研究了Google Scholar、Web of Science、Scopus、Astrophysics Data System(ADS)和INSPIRE等书目数据库在引文合并中的处理方法。Google Scholar和ADS都将两个版本的所有引文合并到出版商版本中,而合并后的引文则累积到INSPIRE存储库中的arXiv版本中。所有这些方法都忽略了arXiv存放版本的类别和相应的可用日期。而Web of Science和Scopus则分别统计两个版本的被引次数,这很可能是将它们视为两篇独立的文章。专注于期刊文章,也出现了arXiv电子打印,我们将它们分为两类,并确定两个公开的文章日期作为引用统计的起点。我们提出了四个可行的方案,以巩固两个版本的文章的引文计数,并提出了一个通用的计划的基础上的研究成果。此外,我们调查了1998年至2018年在arXiv.org上的“计算机科学-数字图书馆”主题(cs.DL)中的2,662份电子印刷品,并手动计算了arXiv存放的文章的合并引用计数以及相应的引用合并方案。此外,这些引文整合方法被应用于文章,作者和期刊的评价。实证检验证明了本文提出的方案的可行性。
Merging the citation counts of arXiv-deposited e-prints (arXiv version) with those of their corresponding published journal articles (publisher version) is an important issue in citation analysis. Using examples of arXiv-deposited e-prints, this article adopts a manual approach to investigate the processing methods used by bibliographic repositories such as Google Scholar, Web of Science, Scopus, Astrophysics Data System (ADS), and INSPIRE for the citation merging. Both Google Scholar and ADS consolidate all citations from the two versions into the publisher one, whereas the consolidated citations are accumulated into the arXiv version in the INSPIRE repository. All these methods ignore the categories of the arXiv-deposited versions and the corresponding availability dates. As for Web of Science and Scopus, they count the citations of the two versions separately, which is likely regarding them as two independent articles. Focusing on journal articles that also appeared as arXiv e-prints, we classify them into two categories and identify two public availability dates of articles as the starting point of citation statistics. We present four feasible schemes to consolidate citation counts for the articles with both versions and also propose a universal scheme based on the research output. Furthermore, we investigated 2,662 e-prints in the "Computer Science-Digital Libraries" subject (cs.DL) from 1998 to 2018 in arXiv.org and manually calculated the consolidated citation counts of arXiv-deposited articles with the corresponding citation merging schemes. Furthermore, these citation consolidation methods are applied to the evaluation of articles, authors, and journals. Such empirical testing proves the feasibility of the schemes proposed in this article.