Dataless Text Classification: A Topic Modeling Approach with Document Manifold

Dataless Text Classification: A Topic Modeling Approach with Document Manifold
复制标题

DOI:
10.1145/3269206.3271671
复制
发表时间:
2018-10
期刊:
Proceedings of the 27th ACM International Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
Ximing Li;C. Li;Jinjin Chi;Jihong Ouyang;Chenliang Li
Ximing Li;C. Li;Jinjin Chi;Jihong Ouyang;Chenliang Li
中科院分区:
其他
文献类型:
--
作者:
Ximing Li;C. Li;Jinjin Chi;Jihong Ouyang;Chenliang Li

文献摘要

被引文献

相似文献

近年来,无数据文本分类受到越来越多的关注。它使用类别的种子字来训练分类器,而不是使用昂贵的标签文档。然而,一小组种子词可能会提供非常有限和嘈杂的监督信息,因为许多文档不包含种子词或仅包含不相关的种子词。在本文中,我们使用文档流形来解决这些问题,假设相邻的文档往往被分配到相同的类别标签。遵循这一思想,我们提出了一种新的拉普拉斯种子词主题模型(LapSWTM)。在LapSWTM中,我们将每个文档建模为隐藏类别主题的混合,每个主题对应一个独特的类别。此外,我们假设相邻文档往往具有相似的类别主题分布。这是通过将流形正则化合并到模型的对数似然函数中,然后最大化该正则化目标来实现的。实验结果表明,我们的LapSWTM算法的性能明显优于现有的无数据文本分类算法,甚至在一定程度上与有监督算法相媲美。更重要的是,当种子词稀缺时,它表现得非常好。
Recently, dataless text classification has attracted increasing attention. It trains a classifier using seed words of categories, rather than labeled documents that are expensive to obtain. However, a small set of seed words may provide very limited and noisy supervision information, because many documents contain no seed words or only irrelevant seed words. In this paper, we address these issues using document manifold, assuming that neighboring documents tend to be assigned to a same category label. Following this idea, we propose a novel Laplacian seed word topic model (LapSWTM). In LapSWTM, we model each document as a mixture of hidden category topics, each of which corresponds to a distinctive category. Also, we assume that neighboring documents tend to have similar category topic distributions. This is achieved by incorporating a manifold regularizer into the log-likelihood function of the model, and then maximizing this regularized objective. Experimental results show that our LapSWTM significantly outperforms the existing dataless text classification algorithms and is even competitive with supervised algorithms to some extent. More importantly, it performs extremely well when the seed words are scarce.