A Pseudo Label based Dataless Naive Bayes Algorithm for Text Classification with Seed Words

A Pseudo Label based Dataless Naive Bayes Algorithm for Text Classification with Seed Words
复制标题

DOI:
--
复制
发表时间:
2018-08
期刊:
--
影响因子:
--
通讯作者:
Ximing Li;Bo Yang
Ximing Li;Bo Yang
中科院分区:
其他
文献类型:
--
作者:
Ximing Li;Bo Yang

文献摘要

被引文献

相似文献

传统的监督文本分类器需要大量的手动标记的文件,这往往是昂贵的获得。最近,无数据文本分类吸引了更多的关注,因为它只需要非常少的种子词的类别是便宜得多。在本文中,我们开发了一个基于伪标签的无数据朴素贝叶斯(PL-DNB)分类器与种子词。我们使用种子词出现为每个文档初始化伪标签,并采用期望最大化算法以半监督的方式训练PL-DNB。使用种子词出现和标签后验的估计的混合迭代地更新伪标签。为了避免噪声伪标签,我们还在伪标签更新步骤中考虑最近邻文档的信息,即,保留文档的局部邻近结构。我们的经验表明,PL-DNB优于传统的无数据文本分类算法与种子词。特别是,PL-DNB在不平衡数据集上表现良好。
Traditional supervised text classifiers require a large number of manually labeled documents, which are often expensive to obtain. Recently, dataless text classification has attracted more attention, since it only requires very few seed words of categories that are much cheaper. In this paper, we develop a pseudo-label based dataless Naive Bayes (PL-DNB) classifier with seed words. We initialize pseudo-labels for each document using seed word occurrences, and employ the expectation maximization algorithm to train PL-DNB in a semi-supervised manner. The pseudo-labels are iteratively updated using a mixture of seed word occurrences and estimations of label posteriors. To avoid noisy pseudo-labels, we also consider the information of nearest neighboring documents in the pseudo-label update step, i.e., preserving local neighborhood structure of documents. We empirically show that PL-DNB outperforms traditional dataless text classification algorithms with seed words. Especially, PL-DNB performs well on the imbalanced dataset.