Co-citation and Co-authorship Networks of Statisticians

Co-citation and Co-authorship Networks of Statisticians
复制标题

DOI:
10.1080/07350015.2021.1978469
复制
发表时间:
2021-11-10
影响因子:
3
通讯作者:
Li, Wanshan
Li, Wanshan
中科院分区:
数学2区
文献类型:
--
作者:
Ji, Pengsheng;Jin, Jiashun;Li, Wanshan

文献摘要

被引文献

相似文献

我们收集并清理了统计出版物的大型数据集。该数据集包含统计、概率和机器学习领域 36 个代表性期刊上发表的 83 篇、331 篇文章的共同作者关系和引用关系,时间跨度为 41 年。该数据集使我们能够构建许多不同的网络,并激发了有关统计界的研究模式和趋势、研究影响和网络拓扑的许多研究问题。在本文中,我们重点关注(i)使用引用关系来估计作者的研究兴趣,以及(ii)使用合著者关系来研究网络拓扑。使用我们构建的共被引网络,我们发现了一个“统计三角形”,让人想起统计哲学三角形(Efron 1998)。我们提出了构建统计学家“研究地图”的新方法,以及给定作者的“研究轨迹”,以可视化他/她的研究兴趣的演变。使用我们构建的共同作者网络,我们发现了一个多层社区树并生成了桑基图来可视化不同子区域的作者迁移。我们还提出了几个衡量个体作者研究多样性的新指标。我们发现“贝叶斯”、“生物统计学”和“非参数”是统计学的三个主要领域。我们还确定了 15 个子领域,每个子领域都可以被视为主要领域的加权平均值,并确定了形成共同作者社区的几个根本原因。我们还发现,在我们研究的 41 年时间窗口中,统计学家的研究兴趣发生了显着变化:某些领域(例如生物统计学、高维数据分析等)变得越来越受欢迎。统计学家的研究多样性可能低于我们的预期。例如,对于大多数作者的个性化网络,所提出的显着性检验的 p 值相对较大。
We collected and cleaned a large dataset on publications in statistics. The dataset consists of the co-author relationships and citation relationships of 83, 331 articles published in 36 representative journals in statistics, probability, and machine learning, spanning 41 years. The dataset allows us to construct many different networks, and motivates a number of research problems about the research patterns and trends, research impacts, and network topology of the statistics community. In this article we focus on (i) using the citation relationships to estimate the research interests of authors, and (ii) using the co-author relationships to study the network topology. Using co-citation networks we constructed, we discover a "statistics triangle," reminiscent of the statistical philosophy triangle (Efron 1998). We propose new approaches to constructing the "research map" of statisticians, as well as the "research trajectory" for a given author to visualize his/her research interest evolvement. Using co-authorship networks we constructed, we discover a multi-layer community tree and produce a Sankey diagram to visualize the author migrations in different sub-areas. We also propose several new metrics for research diversity of individual authors. We find that "Bayes," "Biostatistics," and "Nonparametric" are three primary areas in statistics. We also identify 15 sub-areas, each of which can be viewed as a weighted average of the primary areas, and identify several underlying reasons for the formation of co-authorship communities. We also find that the research interests of statisticians have evolved significantly in the 41-year time window we studied: some areas (e.g., biostatistics, high-dimensional data analysis, etc.) have become increasingly more popular. The research diversity of statisticians may be lower than we might have expected. For example, for the personalized networks of most authors, the p-values of the proposed significance tests are relatively large.