uFACT: Unfaithful Alien-Corpora Training for Semantically Consistent Data-to-Text Generation

uFACT: Unfaithful Alien-Corpora Training for Semantically Consistent Data-to-Text Generation
复制标题

DOI:
10.18653/v1/2022.findings-acl.223
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Tisha Anders;Alexandru Coca;B. Byrne
Tisha Anders;Alexandru Coca;B. Byrne
中科院分区:
其他
文献类型:
--
作者:
Tisha Anders;Alexandru Coca;B. Byrne

文献摘要

相似文献

我们提出了uFACT(Un-Faithful Alien Corpora Training),一种用于数据到文本(d2 t)生成模型的训练语料库构建方法。我们表明,在uFACT数据集上训练的d2 t模型生成的话语比仅在目标语料库上训练的模型更准确地表示数据源的语义内容。我们的方法是增加一个给定的目标语料库的训练集与外国语料库,具有不同的语义表示。我们发现,虽然它是重要的,从目标语料库的忠实数据,其他语料库的忠实性只起着次要的作用。因此,uFACT数据集可以用大量不忠实的数据来构建。我们展示了如何利用uFACT来获得最先进的WebNLG基准测试结果,使用METEOR作为我们的性能指标。此外,我们研究了使用PANEL度量的生成忠实度对训练语料库结构的敏感性,并在WebNLG上为该度量提供了基线(Gardent等人,2017)基准,以便于与未来的工作进行比较。
We propose uFACT (Un-Faithful Alien Corpora Training), a training corpus construction method for data-to-text (d2t) generation models. We show that d2t models trained on uFACT datasets generate utterances which represent the semantic content of the data sources more accurately compared to models trained on the target corpus alone. Our approach is to augment the training set of a given target corpus with alien corpora which have different semantic representations. We show that while it is important to have faithful data from the target corpus, the faithfulness of additional corpora only plays a minor role. Consequently, uFACT datasets can be constructed with large quantities of unfaithful data. We show how uFACT can be leveraged to obtain state-of-the-art results on the WebNLG benchmark using METEOR as our performance metric. Furthermore, we investigate the sensitivity of the generation faithfulness to the training corpus structure using the PARENT metric, and provide a baseline for this metric on the WebNLG (Gardent et al., 2017) benchmark to facilitate comparisons with future work.