Text categorization based on regularized linear classification methods

Text categorization based on regularized linear classification methods
复制标题

DOI:
10.1023/a:1011441423217
复制
发表时间:
2001-01-01
期刊:
INFORMATION RETRIEVAL
影响因子:
--
通讯作者:
Oles, FJ
Oles, FJ
中科院分区:
其他
文献类型:
--
作者:
Zhang, T;Oles, FJ

文献摘要

被引文献

相似文献

线性最小二乘拟合法(LLSF)、Logistic回归和支持向量机等线性分类方法已被应用于文本分类问题。这些方法通过找到近似地将一类文档向量从其补集中分离出来的超平面来共享相似性。然而,到目前为止,支持向量机被认为是特殊的,因为它们已经被证明达到了最先进的性能。因此,值得了解的是,这种良好的性能是支持向量机设计所独有的,还是其他线性分类方法也可以实现的。在本文中,我们比较了一些已知的线性分类方法以及正则化线性系统框架下的一些变体。我们将讨论这些算法的统计和数值特性,重点是文本分类。我们还将提供一些数值实验,以说明这些算法在一些数据集上。
A number of linear classification methods such as the linear least squares fit (LLSF), logistic regression, and support vector machines (SVM's) have been applied to text categorization problems. These methods share the similarity by finding hyperplanes that approximately separate a class of document vectors from its complement. However, support vector machines are so far considered special in that they have been demonstrated to achieve the state of the art performance. It is therefore worthwhile to understand whether such good performance is unique to the SVM design, or if it can also be achieved by other linear classification methods. In this paper, we compare a number of known linear classification methods as well as some variants in the framework of regularized linear systems. We will discuss the statistical and numerical properties of these algorithms, with a focus on text categorization. We will also provide some numerical experiments to illustrate these algorithms on a number of datasets.