An Empirical Study of Contextual Data Augmentation for Japanese Zero Anaphora Resolution

An Empirical Study of Contextual Data Augmentation for Japanese Zero Anaphora Resolution
复制标题

DOI:
10.18653/v1/2020.coling-main.435
复制
发表时间:
2020-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Ryuto Konno;Yuichiroh Matsubayashi;Shun Kiyono;Hiroki Ouchi;Ryo Takahashi;Kentaro Inui
Ryuto Konno;Yuichiroh Matsubayashi;Shun Kiyono;Hiroki Ouchi;Ryo Takahashi;Kentaro Inui
中科院分区:
其他
文献类型:
--
作者:
Ryuto Konno;Yuichiroh Matsubayashi;Shun Kiyono;Hiroki Ouchi;Ryo Takahashi;Kentaro Inui

文献摘要

相似文献

零回指消解(ZAR)的一个关键问题是标记数据的稀缺性。本研究探讨如何有效地缓解这个问题,可以通过数据增强。我们采用了一种最先进的数据增强方法,称为上下文数据增强(CDA),它使用预训练的语言模型生成标记的训练实例。据报道,CDA在其他几个自然语言处理任务中表现良好,包括文本分类和机器翻译。本研究解决了两个未充分探讨的问题,即如何减少数据增强的计算成本,以及如何确保生成的数据的质量。我们还提出了两种使CDA适应ZAR的方法:基于[MASK]的增强和语言控制的掩蔽。因此,在日本ZAR上的实验结果表明,我们的方法有助于提高精度和降低计算成本。我们仔细的分析表明,该方法可以提高增强的训练数据的质量相比,传统的CDA。
One critical issue of zero anaphora resolution (ZAR) is the scarcity of labeled data. This study explores how effectively this problem can be alleviated by data augmentation. We adopt a state-of-the-art data augmentation method, called the contextual data augmentation (CDA), that generates labeled training instances using a pretrained language model. The CDA has been reported to work well for several other natural language processing tasks, including text classification and machine translation. This study addresses two underexplored issues on CDA, that is, how to reduce the computational cost of data augmentation and how to ensure the quality of the generated data. We also propose two methods to adapt CDA to ZAR: [MASK]-based augmentation and linguistically-controlled masking. Consequently, the experimental results on Japanese ZAR show that our methods contribute to both the accuracy gainand the computation cost reduction. Our closer analysis reveals that the proposed method can improve the quality of the augmented training data when compared to the conventional CDA.