Multi-task Learning-based Text Classification with Subword-Phrase Extraction

Multi-task Learning-based Text Classification with Subword-Phrase Extraction
复制标题

DOI:
10.1145/3568562.3568635
复制
发表时间:
2022-12
期刊:
Proceedings of the 11th International Symposium on Information and Communication Technology
影响因子:
--
通讯作者:
Yusuke Kimura;Takahiro Komamizu;K. Hatano
Yusuke Kimura;Takahiro Komamizu;K. Hatano
中科院分区:
其他
文献类型:
--
作者:
Yusuke Kimura;Takahiro Komamizu;K. Hatano

文献摘要

相似文献

基于深度学习的文本分类方法训练了大量的文本,取得了优于传统方法的性能。除了它的成功之外,多任务学习已经成为一种很有前途的文本分类方法;例如,一种多任务学习方法将命名实体识别作为文本分类的辅助任务。现有的基于MTL的文本分类方法依赖于使用监督标签的辅助任务,这需要大量的人力和/或财力来创建。为了减少这些工作量,本文提出了一种基于多任务学习的文本分类框架,该框架减少了有监督的标签创建的额外工作量。实现这一点的一个基本想法是利用由子词组成的短语表达(称为子词-短语)。就我们所知,在子词短语之上还没有文本分类方法,因为子词并不总是表达一组连贯的含义。该框架增加了子词短语识别作为辅助任务,并利用子词短语进行文本分类。为了实现低成本的辅助识别任务,该框架以无监督的方式提取子词短语。在五个常用的文本分类数据集上的实验评估表明,子词短语识别作为辅助任务参与是有效的。它还显示了与最先进的方法的比较结果。
Text classification using deep learning, which is trained with a tremendous amount of text, has achieved superior performance than traditional methods. In addition to its success, multi-task learning has become a promising approach for text classification; for instance, a multi-task learning approach employs named entity recognition as an auxiliary task for text classification. The existing MTL-based text classification methods depend on auxiliary tasks using supervised labels, which require large human and/or financial efforts to create. To reduce these efforts, this paper proposes a multi-task learning-based text classification framework which reduces the additional efforts on supervised label creation. A basic idea to realize this is that to utilize phrasal expressions consisting of subwords (called subword-phrase). To the best of our knowledge, there has been no text classification approach on top of subword-phrases, because subwords do not always express a coherent set of meanings. The proposed framework is new to add subword-phrase recognition as an auxiliary task, and to utilize subword-phrases for text classification. To realize the low-cost auxiliary recognition task, the framework extracts subword-phrases in an unsupervised manner. The experimental evaluation of the five popular datasets for text classification showcases the effectiveness of the involvement of the subword-phrase recognition as an auxiliary task. It also shows comparative results with the state-of-the-art method.