Dataless Text Classification with Descriptive LDA

Dataless Text Classification with Descriptive LDA
复制标题

DOI:
10.1609/aaai.v29i1.9506
复制
发表时间:
2015-01
期刊:
--
影响因子:
--
通讯作者:
Xingyuan Chen;Yunqing Xia;Peng Jin;John A. Carroll
Xingyuan Chen;Yunqing Xia;Peng Jin;John A. Carroll
中科院分区:
其他
文献类型:
--
作者:
Xingyuan Chen;Yunqing Xia;Peng Jin;John A. Carroll

文献摘要

被引文献

相似文献

手动标记文档以训练文本分类器是昂贵且耗时的。此外,在标记的文档上训练的分类器可能遭受过拟合和适应性问题。无数据文本分类(DLTC)已经被提出作为这些问题的解决方案,因为它不需要标记的文档。DLTC先前的研究使用维基百科内容的显式语义分析来测量文档之间的语义距离,这反过来又用于基于最近邻居对测试文档进行分类。基于语义的DLTC方法有一个主要缺点,即它依赖于一个大规模的,精细编译的语义知识库,这是很难获得在许多情况下。在本文中,我们提出了一种新的模型,描述LDA(DescLDA),它执行DLTC只有类别描述词和未标记的文档。在DescLDA中,LDA模型与描述设备组装,以推断Dirichlet先验从先前的描述性文档创建的类别描述词。然后LDA使用Dirichlet先验从未标记的文档中诱导类别感知的潜在主题。在20个Newsgroups和RCV 1数据集上的实验结果表明:(1)我们的DLTC方法比基于语义的DLTC基线方法更有效;(2)我们的DLTC方法的准确率非常接近最先进的监督文本分类方法。由于既不需要外部知识资源,也不需要标记文档,因此我们的DLTC方法适用于更广泛的场景。
Manually labeling documents for training a text classifier is expensive and time-consuming. Moreover, a classifier trained on labeled documents may suffer from overfitting and adaptability problems. Dataless text classification (DLTC) has been proposed as a solution to these problems, since it does not require labeled documents. Previous research in DLTC has used explicit semantic analysis of Wikipedia content to measure semantic distance between documents, which is in turn used to classify test documents based on nearest neighbours. The semantic-based DLTC method has a major drawback in that it relies on a large-scale, finely-compiled semantic knowledge base, which is difficult to obtain in many scenarios. In this paper we propose a novel kind of model, descriptive LDA (DescLDA), which performs DLTC with only category description words and unlabeled documents. In DescLDA, the LDA model is assembled with a describing device to infer Dirichlet priors from prior descriptive documents created with category description words. The Dirichlet priors are then used by LDA to induce category-aware latent topics from unlabeled documents. Experimental results with the 20Newsgroups and RCV1 datasets show that: (1) our DLTC method is more effective than the semantic-based DLTC baseline method; and (2) the accuracy of our DLTC method is very close to state-of-the-art supervised text classification methods. As neither external knowledge resources nor labeled documents are required, our DLTC method is applicable to a wider range of scenarios.