Benefits of Data Augmentation for NMT-based Text Normalization of User-Generated Content

Benefits of Data Augmentation for NMT-based Text Normalization of User-Generated Content
复制标题

数据增强对用户生成内容的基于 NMT 的文本规范化的好处

DOI:
10.18653/v1/d19-5536
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Veronique Hoste
Veronique Hoste
中科院分区:
--
文献类型:
--
作者:
C. Veliz;Orphée De Clercq;Veronique Hoste

文献摘要

被引文献

相似文献

书面用户生成内容(UGC)最持久的特征之一是使用非标准词汇。这一特点增加了自动处理和分析UGC的难度。文本规范化是将词汇变体转换为规范形式的任务,通常用作传统NLP任务的预处理步骤,以克服NLP系统在应用于UGC时遇到的性能下降。在这项工作中,我们遵循神经机器翻译方法来进行文本规范化。为了训练这样的编码器-解码器模型,需要大量的句子对并行训练语料库。然而,使用UGC及其规范化版本获得大型数据集并不容易,特别是对于英语以外的语言。在本文中,我们将探讨如何克服这种数据瓶颈的荷兰语,低资源的语言。我们开始了一个小的公开的并行荷兰数据集,包括三个UGC流派,并比较两种不同的方法。首先是手动规范化和添加训练数据,这是一项既耗时又费钱的任务。第二种方法是一套数据扩充技术,通过将现有资源转换为合成的非标准形式来增加数据大小。我们的研究结果表明,虽然不同的方法产生类似的结果在测试集中的归一化问题,他们也引入了大量的过度归一化。
One of the most persistent characteristics of written user-generated content (UGC) is the use of non-standard words. This characteristic contributes to an increased difficulty to automatically process and analyze UGC. Text normalization is the task of transforming lexical variants to their canonical forms and is often used as a pre-processing step for conventional NLP tasks in order to overcome the performance drop that NLP systems experience when applied to UGC. In this work, we follow a Neural Machine Translation approach to text normalization. To train such an encoder-decoder model, large parallel training corpora of sentence pairs are required. However, obtaining large data sets with UGC and their normalized version is not trivial, especially for languages other than English. In this paper, we explore how to overcome this data bottleneck for Dutch, a low-resource language. We start off with a small publicly available parallel Dutch data set comprising three UGC genres and compare two different approaches. The first is to manually normalize and add training data, a money and time-consuming task. The second approach is a set of data augmentation techniques which increase data size by converting existing resources into synthesized non-standard forms. Our results reveal that, while the different approaches yield similar results regarding the normalization issues in the test set, they also introduce a large amount of over-normalizations.