Coordination Generation via Synchronized Text-Infilling

Coordination Generation via Synchronized Text-Infilling
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Hiroki Teranishi;Yuji Matsumoto
Hiroki Teranishi;Yuji Matsumoto
中科院分区:
其他
文献类型:
--
作者:
Hiroki Teranishi;Yuji Matsumoto

文献摘要

相似文献

从大规模预训练的语言模型中生成用于监督学习的合成数据,提高了多个NLP任务的性能,特别是在低资源场景中。特别是,许多数据增强的研究采用掩蔽语言模型来替换句子中的其他单词。然而,它们中的大多数是在句子分类任务中进行评估的,并且不能立即应用于与句子结构相关的任务。在本文中,我们提出了一个简单而有效的方法来生成句子的坐标结构,其连接词的边界是明确指定的。对于句子中的给定跨度,我们的方法以两种方式(“X和[mask]",“[mask]和X”)嵌入带有协调连词的掩码,并迫使掩码语言模型用相同的文本填充两个空白。为了实现这一点,我们引入了BERT和T5模型的解码方法,并限制了不同掩码的预测是同步的。此外,我们开发了一个训练框架,有效地选择合成的例子监督协调消歧任务。我们证明,我们的方法产生有前途的协调实例,在低资源设置的任务提供收益。
Generating synthetic data for supervised learning from large-scale pre-trained language models has enhanced performances across several NLP tasks, especially in low-resource scenarios. In particular, many studies of data augmentation employ masked language models to replace words with other words in a sentence. However, most of them are evaluated on sentence classification tasks and cannot immediately be applied to tasks related to the sentence structure. In this paper, we propose a simple yet effective approach to generating sentences with a coordinate structure in which the boundaries of its conjuncts are explicitly specified. For a given span in a sentence, our method embeds a mask with a coordinating conjunction in two ways (”X and [mask]”, ”[mask] and X”) and forces masked language models to fill the two blanks with an identical text. To achieve this, we introduce decoding methods for BERT and T5 models with the constraint that predictions for different masks are synchronized. Furthermore, we develop a training framework that effectively selects synthetic examples for the supervised coordination disambiguation task. We demonstrate that our method produces promising coordination instances that provide gains for the task in low-resource settings.