Rewarding Semantic Similarity under Optimized Alignments for AMR-to-Text Generation

Rewarding Semantic Similarity under Optimized Alignments for AMR-to-Text Generation
复制标题

DOI:
10.18653/v1/2022.acl-short.80
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Lisa Jin;D. Gildea
Lisa Jin;D. Gildea
中科院分区:
其他
文献类型:
--
作者:
Lisa Jin;D. Gildea

文献摘要

相似文献

对抗暴露偏差的一种常见方法是将评估指标的分数作为强化学习(RL)的奖励。利用上下文的嵌入似乎比它们的n-gram匹配对应物更灵活,因此是理想的训练奖励。然而,诸如BERTScore之类的度量会使候选令牌和引用令牌greetly对齐,这可以允许系统输出接收相对于引用的超额信用。此外,过去的方法具有语义相似性奖励遭受重复输出和过拟合。我们解决这些问题,提出的指标,取代BERTScore中的贪婪对齐优化的。我们在模型的训练令牌嵌入上计算它们,以防止域不匹配。我们的模型优化离散对齐指标始终优于AMR到文本生成的交叉熵和BLEU奖励基线。此外,我们发现这种方法与非RL设置相比具有稳定的训练。
A common way to combat exposure bias is by applying scores from evaluation metrics as rewards in reinforcement learning (RL). Metrics leveraging contextualized embeddings appear more flexible than their n-gram matching counterparts and thus ideal as training rewards. However, metrics such as BERTScore greedily align candidate and reference tokens, which can allow system outputs to receive excess credit relative to a reference. Furthermore, past approaches featuring semantic similarity rewards suffer from repetitive outputs and overfitting. We address these issues by proposing metrics that replace the greedy alignments in BERTScore with optimized ones. We compute them on a model’s trained token embeddings to prevent domain mismatch. Our model optimizing discrete alignment metrics consistently outperforms cross-entropy and BLEU reward baselines on AMR-to-text generation. In addition, we find that this approach enjoys stable training compared to a non-RL setting.