Text classification based on multi-word with support vector machine

Text classification based on multi-word with support vector machine
复制标题

DOI:
10.1016/j.knosys.2008.03.044
复制
发表时间:
2008-12
期刊:
Knowl. Based Syst.
影响因子:
--
通讯作者:
Wen Zhang;Taketoshi Yoshida;Xijin J. Tang
Wen Zhang;Taketoshi Yoshida;Xijin J. Tang
中科院分区:
其他
文献类型:
--
作者:
Wen Zhang;Taketoshi Yoshida;Xijin J. Tang

文献摘要

被引文献

相似文献

支持文本挖掘的主要主题之一是文本表示,即寻找合适的术语将文档转换为数字向量。近年来,人们在这方面做了大量的工作,利用向量空间模型(VSM)来丰富文本表示,以提高分类、聚类等文本挖掘技术的性能。本文主要关注的是使用多个词来表示文本对分类性能的影响。首先,提出了一种实用的基于句法结构的文档多词抽取方法。其次,提出了两种表示文档的策略,即一般概念表示和子主题表示。特别是,提出了动态k不匹配来确定长多词的存在,该多词是文档内容的子主题。最后,我们分别用多词表示法对路透社-21578文档进行了分类实验。我们以单个词的表征性能为基线,在没有语言预处理的情况下,以特征集维度最大的词作为表征的基线。此外,还对比了支持向量机中的线性核和非线性多项式核进行分类,考察了核类型对分类性能的影响。并以不同的百分比从特征集中剔除信息增益(IG)较低的指标项,以观察各种分类方法的稳健性。实验表明,在多词表示中,一般概念表示的子主题分类性能优于一般概念表示,线性核优于支持向量机的非线性核分类。不同表示策略对分类性能的影响大于不同支持向量机核对分类性能的影响。此外,使用单个单词的表示优于使用多个单词的任何表示。这与大多数人在使用支持向量机进行分类时关于语言预处理对文档特征的作用的观点是一致的。
One of the main themes supporting text mining is text representation, i.e., looking for the appropriate terms to transfer the documents into numerical vectors. Recently, many efforts have been invested on this topic to enrich text representation using vector space model (VSM) to improve the performances of text mining techniques such as classification, clustering, etc. The main concern of this paper is to investigate the effectiveness of using multi-words for text representation on the performances of classification. Firstly, a practical method is proposed to implement the multi-word extraction from documents based on the syntactical structure. Secondly, two strategies as general concept representation and subtopic representation are presented to represent the documents using the extracted multi-words. Especially, the dynamic k-mismatch is proposed to determine the presence of a long multi-word which is a subtopic of the content of a document. Finally, we carried out a series of experiments on classifying the Reuters-21578 documents using the representations with multi-words, respectively. We used the performance of representation in individual words as the baseline, which has the largest dimension of feature set for representation without linguistic preprocessing. Moreover, linear kernel and non-linear polynomial kernel in support vector machines (SVM) are examined comparatively for classification to investigate the effect of kernel type on the performance of classification. And the index terms with low information gain (IG) are removed from the feature set at different percentage to observe the robustness of each classification method. Our experiments demonstrate that in multi-word representation, subtopic of general concept representation outperforms the general concept representation and the linear kernel outperforms non-linear kernel of SVM in classifying the Reuters data. And the effect of applying different representation strategies is greater than the effect of applying the different SVM kernels on classification performance. Furthermore, the representation using individual words outperforms any representation using multi-words. This is consistent with the most opinions concerning the role of linguistic preprocessing on documents’ features when using SVM for classification.