Multi-co-training for document classification using various document representations: TF-IDF, LDA, and Doc2Vec

Multi-co-training for document classification using various document representations: TF-IDF, LDA, and Doc2Vec
复制标题

DOI:
10.1016/j.ins.2018.10.006
复制
发表时间:
2019-03-01
影响因子:
8.1
通讯作者:
Kang, Pilsung
Kang, Pilsung
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kim, Donghwa;Seo, Deokseong;Kang, Pilsung

文献摘要

被引文献

相似文献

文档分类的目的是为指定的文档分配最合适的标签。文档分类的主要挑战是标签信息不足和非结构化稀疏格式。半监督学习(SSL)方法可以有效地解决前一个问题,而考虑多个文档表示方案可以解决后一个问题。协同训练是一种流行的SSL方法,它试图在同一个示例的特征子集方面利用各种视角。在本文中,我们提出了多协同训练(MCT),以提高文档分类的性能。为了增加各种特征集的分类,我们转换文档使用三种文档表示方法:词频逆文档频率(TF-IDF)的词袋计划的基础上,主题分布的基础上潜在的狄利克雷分配(LDA),和基于神经网络的文档嵌入称为文档向量(Doc 2 Vec)。实验结果表明,建议的MCT是鲁棒的参数变化,并优于基准方法在各种条件下。(C)2018爱思唯尔公司All rights reserved.
The purpose of document classification is to assign the most appropriate label to a specified document. The main challenges in document classification are insufficient label information and unstructured sparse format. A semi-supervised learning (SSL) approach could be an effective solution to the former problem, whereas the consideration of multiple document representation schemes can resolve the latter problem. Co-training is a popular SSL method that attempts to exploit various perspectives in terms of feature subsets for the same example. In this paper, we propose multi-co-training (MCT) for improving the performance of document classification. In order to increase the variety of feature sets for classification, we transform a document using three document representation methods: term frequency-inverse document frequency (TF-IDF) based on the bag-of-words scheme, topic distribution based on latent Dirichlet allocation (LDA), and neural-network based document embedding known as document to vector (Doc2Vec). The experimental results demonstrate that the proposed MCT is robust to parameter changes and outperforms benchmark methods under various conditions. (C) 2018 Elsevier Inc. All rights reserved.