Algorithms of nonlinear document clustering based on fuzzy multiset model

Algorithms of nonlinear document clustering based on fuzzy multiset model
复制标题

DOI:
10.1002/int.20263
复制
发表时间:
2008-02
影响因子:
7
通讯作者:
Kiyotaka Mizutani;R. Inokuchi;S. Miyamoto
Kiyotaka Mizutani;R. Inokuchi;S. Miyamoto
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kiyotaka Mizutani;R. Inokuchi;S. Miyamoto

文献摘要

被引文献

相似文献

模糊多集是一种适用于信息检索的模型,因为它具有同时表示元素的个数和属性程度的数学结构。因此,模糊多集也可以作为一种适合于文档聚类的模型。提出了一种基于模糊多集模型的文档聚类算法。将余弦相关的标准邻近度推广到多集模型中,并将两种非线性聚类技术应用到现有的聚类方法中。其中一个引入了用于控制集群卷大小的变量;另一个是支持向量机中使用的内核技巧。此外,还研究了基于竞争学习的聚类算法。当使用核技巧时,高维特征空间中数据的分类配置通过自组织映射被可视化。给出了使用人工数据和真实文档数据的两个数值算例,并讨论了所提方法的效果。©2008威利期刊公司。
Fuzzy multiset is applicable as a model of information retrieval because it has the mathematical structure that expresses the number and the degree of attribution of an element simultaneously. Therefore, fuzzy multisets can be used also as a suitable model for document clustering. This paper aims at developing clustering algorithms based on a fuzzy multiset model for document clustering. The standard proximity measure of the cosine correlation is generalized in the multiset model, and two nonlinear clustering techniques are applied to the existing clustering methods. One introduces a variable for controlling cluster volume sizes; the other one is a kernel trick used in support vector machines. Moreover, clustering by competitive learning is also studied. When the kernel trick has been used the classification configuration of data in a high‐dimensional feature space is visualized by self‐organizing maps. Two numerical examples, which use an artificial data and real document data, are shown and effects of the proposed methods are discussed. © 2008 Wiley Periodicals, Inc.