Feature Selection Using Support Vector Machines

Feature Selection Using Support Vector Machines
复制标题

DOI:
10.2495/data020271
复制
发表时间:
2002-09
期刊:
WIT Transactions on Information and Communication Technologies
影响因子:
--
通讯作者:
J. Brank;M. Grobelnik;Natasa Milic-Frayling;D. Mladenić
J. Brank;M. Grobelnik;Natasa Milic-Frayling;D. Mladenić
中科院分区:
其他
文献类型:
--
作者:
J. Brank;M. Grobelnik;Natasa Milic-Frayling;D. Mladenić

文献摘要

被引文献

相似文献

文本分类是将自然语言文档分类到一组预定义的类别中的任务。文档通常由向量空间模型下的稀疏向量表示,其中词汇表中的每个单词被映射到一个坐标轴,并且其在文档中的出现在表示该文档的向量中产生一个非零分量。当在大量文档集合上训练分类器时,与这些向量的处理相关的时间和存储器要求可能是禁止的。这就要求使用特征选择方法,不仅要减少特征的数量,而且要增加文档向量的稀疏性。提出了一种基于线性支持向量机的特征选择方法。首先,我们在训练数据的一个子集上训练线性SVM,并仅保留那些与分离阳性和阴性样本的所得超平面的法线的高权重分量(在绝对值意义上)相对应的特征。然后,这个减少的特征空间被用来在更大的训练集上训练分类器,因为更多的文档现在可以容纳相同数量的内存。在我们的实验中,我们比较了基于SVM的特征选择的有效性与更传统的特征选择方法,如优势比和信息增益,在实现所需的向量稀疏性和分类性能之间的权衡。实验结果表明,在相同的向量稀疏度下,基于SVM范数的特征选择比基于优势比或基于信息增益的特征选择具有更好的分类性能。
Text categorization is the task of classifying natural language documents into a set of predefined categories. Documents are typically represented by sparse vectors under the vector space model, where each word in the vocabulary is mapped to one coordinate axis and its occurrence in the document gives rise to one nonzero component in the vector representing that document. When training classifiers on large collections of documents, both the time and memory requirements connected with processing of these vectors may be prohibitive. This calls for using a feature selection method, not only to reduce the number of features but also to increase the sparsity of document vectors. We propose a feature selection method based on linear Support Vector Machines (SVMs). First, we train the linear SVM on a subset of training data and retain only those features that correspond to highly weighted components (in absolute value sense) of the normal to the resulting hyperplane that separates positive and negative examples. This reduced feature space is then used to train a classifier over a larger training set because more documents now fit into the same amount of memory. In our experiments we compare the effectiveness of the SVM -based feature selection with that of more traditional feature selection methods, such as odds ratio and information gain, in achieving the desired tradeoff between the vector sparsity and the classification performance. Experimental results indicate that, at the same level of vector sparsity, feature selection based on SVM normals yields better classification performance than odds ratioor information gainbased feature selection when linear SVM classifiers are used.