Training Dynamics for Text Summarization Models

Training Dynamics for Text Summarization Models
复制标题

DOI:
10.18653/v1/2022.findings-acl.163
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Tanya Goyal;Jiacheng Xu;J. Li;Greg Durrett
Tanya Goyal;Jiacheng Xu;J. Li;Greg Durrett
中科院分区:
其他
文献类型:
--
作者:
Tanya Goyal;Jiacheng Xu;J. Li;Greg Durrett

文献摘要

被引文献

相似文献

预训练语言模型(例如 BART)在大型摘要数据集上进行微调时显示出令人印象深刻的结果。然而,人们对这种微调过程知之甚少,包括从预训练中保留了哪些知识,或者如何在迭代中学习内容选择和生成策略。在这项工作中,我们分析了生成模型的训练动态,重点是总结。在不同的数据集(CNN/DM、XSum、MediaSum)和摘要属性(例如抽象性和幻觉)中,我们研究模型在微调过程的不同阶段学到了什么。我们发现,在训练过程的早期,在所有研究的数据集中一致地学习到了复制输入的倾向。另一方面,事实错误,例如对不受支持的事实的幻觉,是在后期阶段习得的,尽管这种行为在不同领域的情况更加不同。基于这些观察,我们探索了修改训练的补充方法:首先,忽略难以学习的高损失标记,其次,忽略在训练过程的后期阶段快速学习的低损失标记。我们表明,这些简单的训练修改允许我们配置模型以实现不同的目标,例如提高事实性或提高抽象性。
Pre-trained language models (e.g. BART) have shown impressive results when fine-tuned on large summarization datasets. However, little is understood about this fine-tuning process, including what knowledge is retained from pre-training time or how content selection and generation strategies are learnt across iterations. In this work, we analyze the training dynamics for generation models, focusing on summarization. Across different datasets (CNN/DM, XSum, MediaSum) and summary properties, such as abstractiveness and hallucination, we study what the model learns at different stages of its fine-tuning process. We find that a propensity to copy the input is learned early in the training process consistently across all datasets studied. On the other hand, factual errors, such as hallucination of unsupported facts, are learnt in the later stages, though this behavior is more varied across domains. Based on these observations, we explore complementary approaches for modifying training: first, disregarding high-loss tokens that are challenging to learn and second, disregarding low-loss tokens that are learnt very quickly in the latter stages of the training process. We show that these simple training modifications allow us to configure our model to achieve different goals, such as improving factuality or improving abstractiveness.