Time-aware Prompting for Text Generation

Time-aware Prompting for Text Generation
复制标题

DOI:
10.48550/arxiv.2211.02162
复制
发表时间:
2022-11
期刊:
--
影响因子:
--
通讯作者:
Shuyang Cao;Lu Wang
Shuyang Cao;Lu Wang
中科院分区:
其他
文献类型:
--
作者:
Shuyang Cao;Lu Wang

文献摘要

相似文献

在本文中,我们研究了将时间戳(例如文档创建日期)合并到生成系统中的效果。研究了两种类型的时间感知提示:(1)以自然语言句子编码文档时间戳的文本提示; (2) 将时间戳转换为连续向量的线性提示。为了探索对未来数据点的外推,我们进一步引入了一个新的数据到文本生成数据集 TempWikiBio,其中包含超过 400 万条按时间顺序排列的英语维基百科传记文章修订版,每条都与结构化个人资料配对。通过 TempWikiBio 上的数据到文本生成、内容传输数据集上的文本到文本生成以及 XSum 上的摘要,我们表明编码器上的线性提示和文本提示提高了所有数据集的生成质量。根据人类评估和敏感性分析,尽管在对稍后提取的数据进行测试时性能下降较少,但线性提示更关注非时间信息,并且对给定时间戳不太敏感。同时,文本提示建立给定时间戳和输出日期之间的关联,从而在输出中产生更多事实时间信息。
In this paper, we study the effects of incorporating timestamps, such as document creation dates, into generation systems. Two types of time-aware prompts are investigated: (1) textual prompts that encode document timestamps in natural language sentences; and (2) linear prompts that convert timestamps into continuous vectors. To explore extrapolation to future data points, we further introduce a new data-to-text generation dataset, TempWikiBio, containing more than 4 millions of chronologically ordered revisions of biographical articles from English Wikipedia, each paired with structured personal profiles. Through data-to-text generation on TempWikiBio, text-to-text generation on the content transfer dataset, and summarization on XSum, we show that linear prompts on encoder and textual prompts improve the generation quality on all datasets. Despite having less performance drop when testing on data drawn from a later time, linear prompts focus more on non-temporal information and are less sensitive to the given timestamps, according to human evaluations and sensitivity analyses. Meanwhile, textual prompts establish the association between the given timestamps and the output dates, yielding more factual temporal information in the output.