Text classification based on multi-word with support vector machine
Text classification based on multi-word with support vector machine
复制标题
DOI:
10.1016/j.knosys.2008.03.044
复制
发表时间:
2008-12
期刊:
影响因子:
--
通讯作者:
Wen Zhang;Taketoshi Yoshida;Xijin J. Tang
中科院分区:
文献类型:
--
作者:
Wen Zhang;Taketoshi Yoshida;Xijin J. Tang
One of the main themes supporting text mining is text representation, i.e., looking for the appropriate terms to transfer the documents into numerical vectors. Recently, many efforts have been invested on this topic to enrich text representation using vector space model (VSM) to improve the performances of text mining techniques such as classification, clustering, etc. The main concern of this paper is to investigate the effectiveness of using multi-words for text representation on the performances of classification. Firstly, a practical method is proposed to implement the multi-word extraction from documents based on the syntactical structure. Secondly, two strategies as general concept representation and subtopic representation are presented to represent the documents using the extracted multi-words. Especially, the dynamic k-mismatch is proposed to determine the presence of a long multi-word which is a subtopic of the content of a document. Finally, we carried out a series of experiments on classifying the Reuters-21578 documents using the representations with multi-words, respectively. We used the performance of representation in individual words as the baseline, which has the largest dimension of feature set for representation without linguistic preprocessing. Moreover, linear kernel and non-linear polynomial kernel in support vector machines (SVM) are examined comparatively for classification to investigate the effect of kernel type on the performance of classification. And the index terms with low information gain (IG) are removed from the feature set at different percentage to observe the robustness of each classification method. Our experiments demonstrate that in multi-word representation, subtopic of general concept representation outperforms the general concept representation and the linear kernel outperforms non-linear kernel of SVM in classifying the Reuters data. And the effect of applying different representation strategies is greater than the effect of applying the different SVM kernels on classification performance. Furthermore, the representation using individual words outperforms any representation using multi-words. This is consistent with the most opinions concerning the role of linguistic preprocessing on documents’ features when using SVM for classification.