Using back-and-forth translation to create artificial augmented textual data for sentiment analysis models

Using back-and-forth translation to create artificial augmented textual data for sentiment analysis models
复制标题

使用来回翻译为情感分析模型创建人工增强文本数据

DOI:
10.1016/j.eswa.2021.115033
复制
发表时间:
2021
影响因子:
8.5
通讯作者:
Zhong Ning
Zhong Ning
中科院分区:
计算机科学1区
文献类型:
--
作者:
Body Thomas;Tao Xiaohui;Li Yuefeng;Li Lin;Zhong Ning

文献摘要

相似文献

使用神经网络训练的情感分析分类模型需要大量数据,但收集这些数据集需要大量时间和资源。尽管人工数据已成功应用于计算机视觉,但用于创建人工增强文本数据的有效且可推广的方法却很少。在本文中,提出了一种基于文本的数据增强方法,称为来回翻译,可用于人为地增加任何自然语言数据集的大小。通过创建增强文本数据并将其添加到原始数据集中,实证实验证明,来回翻译数据增强可以将二元情感分类模型的错误率降低高达 3.4%。这些结果显示出统计显着性。
Sentiment analysis classification models trained using neural networks require large amounts of data, but collecting these datasets requires significant time and resources. Although artificial data has been used successfully in computer vision, there are few effective and generalizable methods for creating artificial augmented text data. In this paper, a text based data augmentation method is proposed called back-and-forth translation that can be used to artificially increase the size of any natural language dataset. By creating augmented text data and adding it to the original dataset, it is demonstrated by empirical experiments that back-and-forth translation data augmentation can reduce the error rate in binary sentiment classification models by up to 3.4%. These results are shown to be statistically significant.