Minimally-Supervised Structure-Rich Text Categorization via Learning on Text-Rich Networks

Minimally-Supervised Structure-Rich Text Categorization via Learning on Text-Rich Networks
复制标题

DOI:
10.1145/3442381.3450114
复制
发表时间:
2021-02
期刊:
Proceedings of the Web Conference 2021
影响因子:
--
通讯作者:
Xinyang Zhang;Chenwei Zhang;Xin Dong;Jingbo Shang;Jiawei Han
Xinyang Zhang;Chenwei Zhang;Xin Dong;Jingbo Shang;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Xinyang Zhang;Chenwei Zhang;Xin Dong;Jingbo Shang;Jiawei Han

文献摘要

相似文献

文本分类是Web内容分析中的一项重要任务。考虑到不断发展的Web数据和新出现的类别,在本文中,我们关注的不是繁琐的监督设置,而是旨在有效地对文档进行分类的最小监督设置,每个类别都标注了几个种子文档。我们认识到,从Web收集的文本往往是结构丰富的,即伴随着各种元数据。人们可以很容易地将语料库组织到一个富文本网络中,将原始文本文档与文档属性、高质量短语、标签表面名称作为节点以及它们的关联作为边进行连接。这样的网络提供了对语料库的不同数据源的整体看法,并使基于网络的分析和深度文本模型训练能够联合优化。因此,我们提出了一种基于富文本网络学习的最小监督分类框架。具体地说,我们联合训练了两个具有不同归纳偏见的模块-一个用于文本理解的文本分析模块和一个用于班级区分、可扩展的网络学习的网络学习模块。每个模块从未标记的文档集生成伪训练标签,并且两个模块通过使用汇集的伪标签来共同训练来相互增强。我们在两个真实世界的数据集上测试了我们的模型。在683个类别的具有挑战性的电子商务产品分类数据集上,我们的实验表明,在每个类别只有三个种子文档的情况下,我们的框架可以达到约92%的准确率,显著优于所有比较的方法;我们的准确率与在大约50K标签文档上训练的有监督的BERT模型仅相差不到2%。
Text categorization is an essential task in Web content analysis. Considering the ever-evolving Web data and new emerging categories, instead of the laborious supervised setting, in this paper, we focus on the minimally-supervised setting that aims to categorize documents effectively, with a couple of seed documents annotated per category. We recognize that texts collected from the Web are often structure-rich, i.e., accompanied by various metadata. One can easily organize the corpus into a text-rich network, joining raw text documents with document attributes, high-quality phrases, label surface names as nodes, and their associations as edges. Such a network provides a holistic view of the corpus’ heterogeneous data sources and enables a joint optimization for network-based analysis and deep textual model training. We therefore propose a novel framework for minimally supervised categorization by learning from the text-rich network. Specifically, we jointly train two modules with different inductive biases – a text analysis module for text understanding and a network learning module for class-discriminative, scalable network learning. Each module generates pseudo training labels from the unlabeled document set, and both modules mutually enhance each other by co-training using pooled pseudo labels. We test our model on two real-world datasets. On the challenging e-commerce product categorization dataset with 683 categories, our experiments show that given only three seed documents per category, our framework can achieve an accuracy of about 92%, significantly outperforming all compared methods; our accuracy is only less than 2% away from the supervised BERT model trained on about 50K labeled documents.