Text categorization with support vector machines.: How to represent texts in input space?

Text categorization with support vector machines.: How to represent texts in input space?
复制标题

DOI:
10.1023/a:1012491419635
复制
发表时间:
2002-01-01
期刊:
影响因子:
7.5
通讯作者:
Kindermann, J
Kindermann, J
中科院分区:
计算机科学3区
文献类型:
--
作者:
Leopold, E;Kindermann, J

文献摘要

被引文献

相似文献

核函数的选择对于支持向量机的大多数应用都是至关重要的。然而,在本文中,我们表明,在文本分类的情况下,词频变换对SVM的性能比内核本身有更大的影响。我们讨论的重要性权重的作用(例如,文件的频率和冗余),这是尚未完全理解的模型的复杂性和计算成本,我们表明,即使在分类一个高度曲折的语言,如德国,可以避免耗时的词形还原或词干。
The choice of the kernel function is crucial to most applications of support vector machines. In this paper, however, we show that in the case of text classification, term-frequency transformations have a larger impact on the performance of SVM than the kernel itself. We discuss the role of importance-weights (e.g. document frequency and redundancy), which is not yet fully understood in the light of model complexity and calculation cost, and we show that time consuming lemmatization or stemming can be avoided even when classifying a highly inflectional language like German.