Text Classification by Combining Different Distance Functions with Weights

Text Classification by Combining Different Distance Functions with Weights
复制标题

DOI:
10.1541/ieejeiss.127.2077
复制
发表时间:
2007
影响因子:
--
通讯作者:
T. Yamada;Naohiro Ishii;Toyoshiro Nakashima
T. Yamada;Naohiro Ishii;Toyoshiro Nakashima
中科院分区:
--
文献类型:
--
作者:
T. Yamada;Naohiro Ishii;Toyoshiro Nakashima

文献摘要

相似文献

文本分类是数据挖掘中的一个重要课题。对于文本分类,目前已经发展了几种方法,如最近邻分析、潜在语义分析等。k-最近邻(kNN)分类是一种众所周知的简单有效的数据分类方法。在kNN的使用中,距离函数对于测量数据之间的距离和相似度非常重要。为了提高kNN分类器的性能,本文提出了一种组合多个距离函数的新方法。利用遗传算法计算距离函数中各元素的权重因子,以保证测量的有效性。为提高分类精度,提出了一种集成处理方法。最后,通过实验表明,本文提出的方法在文本分类中是有效的。
The text classification is an important subject in the data mining. For the text classification, several methods have been developed up to now, as the nearest neighbor analysis, the latent semantic analysis, etc. The k-nearest neighbor (kNN) classification is a well-known simple and effective method for the classification of data in many domains. In the use of the kNN, the distance function is important to measure the distance and the similarity between data. To improve the performance of the classifier by the kNN, a new approach to combine multiple distance functions is proposed here. The weighting factors of elements in the distance function, are computed by GA for the effectiveness of the measurement. Further, an ensemble processing was developed for the improvement of the classification accuracy. Finally, it is shown by experiments that the methods, developed here, are effective in the text classification.