Multi-label dataless text classification with topic modeling

Multi-label dataless text classification with topic modeling
复制标题

具有主题建模的多标签无数据文本分类

DOI:
10.1007/s10115-018-1280-0
复制
发表时间:
2017-11
期刊:
Knowledge and Information Systems (CCF-B类)
影响因子:
--
通讯作者:
Li Chenliang
Li Chenliang
中科院分区:
其他
文献类型:
--
作者:
Zha Daochen;Li Chenliang

文献摘要

参考文献

被引文献

相似文献

手动标记文档既繁琐又昂贵,但这对于训练传统的文本分类器是必不可少的。近年来,人们提出了一些新的文本分类技术来解决这个问题。然而,现有的工作主要集中在单标签分类问题上,即每个文档被限制在属于单一类别。在本文中,我们提出了一个新的种子导向的多标签topicmodel,命名为SMTM。使用与每个类别相关的几个种子词,SMTM对没有任何标记文档的文档集合进行多标签分类。在SMTM中,每个类别都与一个类别主题相关联,该主题涵盖了类别的含义。为了适应多标签文档,我们通过使用spike和slab先验以及弱平滑先验来显式地建模SMTM中的类别稀疏性。也就是说,在不使用任何阈值调优的情况下,SMTM自动为每个文档选择相关的类别。为了结合对种子词的监督,我们提出了一个种子引导的有偏GPU(即广义Pólya urn)采样过程来指导SMTM的主题推理。在两个公共数据集上的实验表明,SMTM比最先进的替代方案具有更好的分类精度,甚至在某些情况下优于监督解决方案。
Manually labeling documents is tedious and expensive, but it is essential for training a traditional text classifier. In recent years, a fewdataless text classificationtechniques have been proposed to address this problem. However, existing works mainly center on single-label classification problems, that is, each document is restricted to belonging to a single category. In this paper, we propose a novelSeed-guidedMulti-labelTopicModel, named SMTM. With a few seed words relevant to each category, SMTM conducts multi-label classification for a collection of documents without any labeled document. In SMTM, each category is associated with a single category-topic which covers the meaning of the category. To accommodate with multi-label documents, we explicitly model the category sparsity in SMTM by usingspike and slabprior and weak smoothing prior. That is, without using any threshold tuning, SMTM automatically selects the relevant categories for each document. To incorporate the supervision of the seed words, we propose a seed-guided biased GPU (i.e., generalized Pólya urn) sampling procedure to guide the topic inference of SMTM. Experiments on two public datasets show that SMTM achieves better classification accuracy than state-of-the-art alternatives and even outperforms supervised solutions in some scenarios.
DOI: 10.1111/j.1467-985x.2009.00614_13.x
发表时间: 2009-10
影响因子: 2
作者:
J. Haigh
通讯作者: J. Haigh
DOI: --
发表时间: 2018-08
期刊: --
影响因子: --
作者:
Ximing Li;Bo Yang
通讯作者: Ximing Li;Bo Yang
DOI: 10.1145/3269206.3271671
发表时间: 2018-10
期刊: Proceedings of the 27th ACM International Conference on Information and Knowledge Management
影响因子: --
作者:
Ximing Li;C. Li;Jinjin Chi;Jihong Ouyang;Chenliang Li
通讯作者: Ximing Li;C. Li;Jinjin Chi;Jihong Ouyang;Chenliang Li
DOI: 10.1609/aaai.v29i1.9506
发表时间: 2015-01
期刊: --
影响因子: --
作者:
Xingyuan Chen;Yunqing Xia;Peng Jin;John A. Carroll
通讯作者: Xingyuan Chen;Yunqing Xia;Peng Jin;John A. Carroll
DOI: 10.1109/asonam.2009.19
发表时间: 2009-07
期刊: 2009 International Conference on Advances in Social Network Analysis and Mining
影响因子: --
作者:
A. Zubiaga;A. P. García-Plaza;Víctor Fresno-Fernández;Raquel Martínez-Unanue
通讯作者: A. Zubiaga;A. P. García-Plaza;Víctor Fresno-Fernández;Raquel Martínez-Unanue