Text classification using string kernels

Text classification using string kernels
复制标题

DOI:
10.1162/153244302760200687
复制
发表时间:
2002-06-01
影响因子:
6
通讯作者:
Watkins, C
Watkins, C
中科院分区:
计算机科学3区
文献类型:
--
作者:
Lodhi, H;Saunders, C;Watkins, C

文献摘要

被引文献

相似文献

我们提出了一种基于使用特殊内核对文本文档进行分类的新颖方法。内核是由所有长度为 k 的子序列生成的特征空间中的内积。子序列是文本中出现的 k 个字符的任意有序序列,但不一定是连续的。子序列按文本中全长的指数衰减因子进行加权,因此强调那些接近连续的出现。即使对于适中的 k 值,直接计算该特征向量也将涉及大量计算,因为特征空间的维度随 k 呈指数增长。尽管存在这一事实,本文仍描述了如何通过动态编程技术有效地评估内积。
We propose a novel approach for categorizing text documents based on the use of a special kernel. The kernel is an inner product in the feature space generated by all subsequences of length k. A subsequence is any ordered sequence of k characters occurring in the text though not necessarily contiguously. The subsequences are weighted by an exponentially decaying factor of their full length in the text, hence emphasising those occurrences that are close to contiguous. A direct computation of this feature vector would involve a prohibitive amount of computation even for modest values of k, since the dimension of the feature space grows exponentially with k. The paper describes how despite this fact the inner product can be efficiently evaluated by a dynamic programming technique.