Hybrid Clustering of Text Mining and Bibliometrics Applied to Journal Sets

Hybrid Clustering of Text Mining and Bibliometrics Applied to Journal Sets
复制标题

DOI:
10.1137/1.9781611972795.5
复制
发表时间:
2009-06
期刊:
--
影响因子:
--
通讯作者:
Xinhai Liu;Shi Yu;Y. Moreau;B. Moor;W. Glänzel;Frizo A. L. Janssens
Xinhai Liu;Shi Yu;Y. Moreau;B. Moor;W. Glänzel;Frizo A. L. Janssens
中科院分区:
其他
文献类型:
--
作者:
Xinhai Liu;Shi Yu;Y. Moreau;B. Moor;W. Glänzel;Frizo A. L. Janssens

文献摘要

被引文献

相似文献

为了获得文本挖掘和文献计量学中相关和互补的信息,将文本内容和引文信息结合起来的混合聚类已经成为一种流行的策略。在本文中,我们提出了一个新的计算框架,结合文本挖掘和文献计量学提供一个映射的期刊集。本文采用了两种不同的混合聚类方法。第一类是集成聚类,它将从单个数据中获得的不同聚类结果组合成一个统一的聚类结果。第二类是核融合,它将异构数据集映射到核空间,并结合核矩阵进行聚类。核可以平均组合,也可以通过优化的加权线性组合模型组合。在本文中,我们提出了一种新的自适应核K-均值聚类算法,结合联合收割机的文本内容和引文信息进行聚类。对2002-2006年出版的1869种期刊进行聚类分析,将该算法与其他方法进行了系统的比较。实验结果表明,基于多个验证指标,我们的混合聚类策略能够提供聚类结果以及最佳的个人数据源。
To obtain correlated and complementary information contained in text mining and bibliometrics, hybrid clustering to incorporate textual content and citation information has become a popular strategy. In this paper, we propose a new computational framework of integrating text mining and bibliometrics to provide a mapping of journal sets. Two different approaches of hybrid clustering methods are applied in this paper. The first category is ensemble clustering, which combines different clustering results obtained from individual data into a consolidated clustering result. The second category is kernel fusion, which maps heterogeneous data sets into the kernel space and combines the kernel matrices for clustering. Kernels can be combined either averagely, or by an optimized weighted linear combination model. In this paper, we propose a novel adaptive kernel K-means clustering algorithm to combine textual content and citation information for clustering. The proposed algorithm is systematically compared with other methods on a clustering problem of 1869 journals published in 2002-2006. Based on several validation indices, the experimental results demonstrate that our hybrid clustering strategy is able to provide clustering result as well as the best individual data source.