Latent Support Measure Machines for Bag-of-Words Data Classification

Latent Support Measure Machines for Bag-of-Words Data Classification
复制标题

DOI:
--
复制
发表时间:
2014-12
期刊:
--
影响因子:
--
通讯作者:
Yuya Yoshikawa;Tomoharu Iwata;H. Sawada
Yuya Yoshikawa;Tomoharu Iwata;H. Sawada
中科院分区:
其他
文献类型:
--
作者:
Yuya Yoshikawa;Tomoharu Iwata;H. Sawada

文献摘要

被引文献

相似文献

在许多分类问题中,输入表示为一组特征,例如,文档的词袋(Bag-of-Words,BoW)表示。支持向量机(SVM)是广泛使用的工具,这样的分类问题。支持向量机的性能通常取决于是否可以正确定义数据点之间的核值。然而,支持向量机的BoW表示有一个主要的弱点,不同的,但语义相似的词的同现不能反映在内核计算。为了克服这一弱点,我们提出了一个基于核的判别分类器的BoW数据,我们称之为潜在的支持度量机(潜在SMM)。利用潜在SMM,潜在向量与每个词汇项相关联,并且每个文档被表示为文档中出现的单词的潜在向量的分布。为了有效地表示分布,我们使用了分布的核嵌入,它包含了分布的高阶矩信息。然后,潜在的SMM找到一个分离的超平面,最大化不同类别的分布之间的边缘,同时估计潜在向量的单词,以提高分类性能。在实验中,我们表明,潜在的SMM达到了最先进的准确性BoW文本分类,是强大的相对于自己的超参数,是有用的可视化的话。
In many classification problems, the input is represented as a set of features, e.g., the bag-of-words (BoW) representation of documents. Support vector machines (SVMs) are widely used tools for such classification problems. The performance of the SVMs is generally determined by whether kernel values between data points can be defined properly. However, SVMs for BoW representations have a major weakness in that the co-occurrence of different but semantically similar words cannot be reflected in the kernel calculation. To overcome the weakness, we propose a kernel-based discriminative classifier for BoW data, which we call the latent support measure machine (latent SMM). With the latent SMM, a latent vector is associated with each vocabulary term, and each document is represented as a distribution of the latent vectors for words appearing in the document. To represent the distributions efficiently, we use the kernel embeddings of distributions that hold high order moment information about distributions. Then the latent SMM finds a separating hyperplane that maximizes the margins between distributions of different classes while estimating latent vectors for words to improve the classification performance. In the experiments, we show that the latent SMM achieves state-of-the-art accuracy for BoW text classification, is robust with respect to its own hyper-parameters, and is useful to visualize words.