Stochastic Tokenization with a Language Model for Neural Text Classification

Stochastic Tokenization with a Language Model for Neural Text Classification
复制标题

DOI:
10.18653/v1/p19-1158
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Tatsuya Hiraoka;Hiroyuki Shindo;Yuji Matsumoto
Tatsuya Hiraoka;Hiroyuki Shindo;Yuji Matsumoto
中科院分区:
其他
文献类型:
--
作者:
Tatsuya Hiraoka;Hiroyuki Shindo;Yuji Matsumoto

文献摘要

相似文献

对于未切分的语言,如日语和汉语,句子的标记化对文本分类的性能有很大影响。句子通常由词法分析器或字节对编码用词或子词分割,然后用词(或子词)表示进行神经网络编码。然而,分割可能是模棱两可的,并且还不清楚分割的令牌对于目标任务是否达到最佳性能。在本文中,我们提出了一种同时学习标记化和文本分类的方法来解决这些问题。我们的模型将用于无监督标记化的语言模型结合到文本分类器中,然后同时训练这两个模型。在训练过程中,我们对每个句子的切分进行了随机采样,从而提高了文本分类的性能。作为文本分类任务,我们对情感分析进行了实验,实验结果表明,该方法比以往的方法取得了更好的性能。
For unsegmented languages such as Japanese and Chinese, tokenization of a sentence has a significant impact on the performance of text classification. Sentences are usually segmented with words or subwords by a morphological analyzer or byte pair encoding and then encoded with word (or subword) representations for neural networks. However, segmentation is potentially ambiguous, and it is unclear whether the segmented tokens achieve the best performance for the target task. In this paper, we propose a method to simultaneously learn tokenization and text classification to address these problems. Our model incorporates a language model for unsupervised tokenization into a text classifier and then trains both models simultaneously. To make the model robust against infrequent tokens, we sampled segmentation for each sentence stochastically during training, which resulted in improved performance of text classification. We conducted experiments on sentiment analysis as a text classification task and show that our method achieves better performance than previous methods.