Concepts of the cover coefficient-based clustering methodology

Concepts of the cover coefficient-based clustering methodology
复制标题

基于覆盖系数的聚类方法的概念

DOI:
10.1145/253495.253526
复制
发表时间:
1985
期刊:
--
影响因子:
--
通讯作者:
E. Ozkarahan
E. Ozkarahan
中科院分区:
--
文献类型:
--
作者:
F. Can;E. Ozkarahan

文献摘要

被引文献

相似文献

文档聚类有几个未解决的问题。其中包括时间和空间复杂度高、相似度阈值确定困难、顺序依赖性、簇中文档分布不均匀以及各种簇启动器确定的任意性。为了在一定程度上克服这些问题,引入了基于覆盖系数的聚类方法。该方法中使用的概念创建了某些新概念、关系和度量,例如索引对聚类的影响、索引的最佳词汇生成以及新的匹配函数。对这些新概念进行了讨论。还包括显示聚类方法和匹配功能有效性的性能实验结果。在这些实验中,还观察到在搜索中获得的大多数文档集中在包含数据库中低百分比文档的少数簇中。
Document clustering has several unresolved problems. Among them are high time and space complexity, difficulty of determining similarity thresholds, order dependence, nonuniform document distribution in clusters, and arbitrariness in determination of various cluster intiators. To overcome these problems to some degree, the cover coefficient based clustering methodology has been introduced. The concepts used in this methodology have created certain new concepts, relationships, and measures such as the effect of indexing on clustering, an optimal vocabulary generation for indexing, and a new matching function. These new concepts are discussed. The result of performance experiments that show the effectiveness of the clustering methodology and the matching function are also included. In these experiments, it has been also observed that the majority of the documents obtained in a search are concentrated in a few clusters containing a low percentage of documents of the database.