Co-occurrence matrices and their applications in information science: Extending ACA to the Web environment

Co-occurrence matrices and their applications in information science: Extending ACA to the Web environment
复制标题

DOI:
10.1002/asi.20335
复制
发表时间:
2006-10-01
影响因子:
--
通讯作者:
Vaughan, Liwen
Vaughan, Liwen
中科院分区:
其他
文献类型:
--
作者:
Leydesdorff, Loet;Vaughan, Liwen

文献摘要

被引文献

相似文献

共现矩阵在信息科学中得到了广泛的应用,如同源矩阵、共词矩阵和共生矩阵。然而,混乱和争议阻碍了对这些数据的适当统计分析。我们认为,根本问题涉及理解各种类型的矩阵的性质。本文讨论了对称引文矩阵和非对称引文矩阵之间的区别,以及可以分别应用于这两个矩阵的适当的统计技术。相似性度量(如皮尔逊相关系数或余弦)不应应用于对称引文矩阵,而可应用于非对称引文矩阵以得出接近度矩阵。文中举例说明了这一论点。然后,该研究将共现矩阵的应用扩展到Web环境中,在Web环境中,可用数据的性质以及数据收集方法不同于传统数据库,如科学引文索引。使用传统的多元分析方法和基于社会网络分析和图论的新可视化软件Pajek,对谷歌学者搜索引擎收集的一组数据进行了分析。
Co-occurrence matrices, such as cocitation, coword, and colink matrices, have been used widely in the information sciences. However, confusion and controversy have hindered the proper statistical analysis of these data. The underlying problem, in our opinion, involved understanding the nature of various types of matrices. This article discusses the difference between a symmetrical cocitation matrix and an asymmetrical citation matrix as well as the appropriate statistical techniques that can be applied to each of these matrices, respectively. Similarity measures (such as the Pearson correlation coefficient or the cosine) should not be applied to the symmetrical cocitation matrix but can be applied to the asymmetrical citation matrix to derive the proximity matrix. The argument is illustrated with examples. The study then extends the application of co-occurrence matrices to the Web environment, in which the nature of the available data and thus data collection methods are different from those of traditional databases such as the Science Citation Index. A set of data collected with the Google Scholar search engine is analyzed by using both the traditional methods of multivariate analysis and the new visualization software Pajek, which is based on social network analysis and graph theory.