Documents Clustering Based on Optimized Compressibility Vector Space

Documents Clustering Based on Optimized Compressibility Vector Space
复制标题

DOI:
10.1109/cise.2009.5363976
复制
发表时间:
2009-12
期刊:
2009 International Conference on Computational Intelligence and Software Engineering
影响因子:
--
通讯作者:
Nuo Zhang;Toshinori Watanabe
Nuo Zhang;Toshinori Watanabe
中科院分区:
其他
文献类型:
--
作者:
Nuo Zhang;Toshinori Watanabe

文献摘要

相似文献

高性能的计算机硬件和宽带接入网络使大规模电子文档的存取和存储成为可能。为了正确处理这些不断增加的文档数量,一个有效的文档表示模型和分类算法一样重要。文本表示方法,如词袋模型和N-gram模型,已被广泛使用。另一种表示方法,称为模式表示方案使用数据压缩(PRDC)最近已被提出。它不仅可以独立处理语言文本数据,而且可以有效地处理多媒体数据。在这项研究中,我们将提出一种方法来改进PRDC方法,并将其与上述两种方法进行比较。性能将在聚类能力方面进行比较。实验结果表明,该方法可以提供更好的性能比其他两种方法,也PRDC。
To access and store large-scale electrical documents becomes possible due to the high performance of computer hardware and broadband accessible network. In order to handle these increasing number of documents properly, a efficient doc- ument representation model is as important as the classification algorithms. Several text representation methods, such as bag- of-words and N-gram models, have been widely used. Another representation approach named pattern representation scheme using data compression (PRDC) has been proposed lately. It does not only independently process data of linguistic text, but also processes multimedia data effectively. In this study, we will propose a method to improve PRDC approach and compare it with the two aforementioned methods. The performances will be compared in terms of clustering ability. Experiment results will show that the proposed method can provide better performance than that of the other two methods and also the PRDC.