ToTTo: A Controlled Table-To-Text Generation Dataset

ToTTo: A Controlled Table-To-Text Generation Dataset
复制标题

DOI:
10.18653/v1/2020.emnlp-main.89
复制
发表时间:
2020-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Ankur P. Parikh;Xuezhi Wang;Sebastian Gehrmann;Manaal Faruqui;Bhuwan Dhingra;Diyi Yang;Dipanjan Das
Ankur P. Parikh;Xuezhi Wang;Sebastian Gehrmann;Manaal Faruqui;Bhuwan Dhingra;Diyi Yang;Dipanjan Das
中科院分区:
其他
文献类型:
--
作者:
Ankur P. Parikh;Xuezhi Wang;Sebastian Gehrmann;Manaal Faruqui;Bhuwan Dhingra;Diyi Yang;Dipanjan Das

文献摘要

被引文献

相似文献

我们提出了ToTTo,一个开放领域的英语表格到文本数据集,拥有超过120,000个训练示例,提出了一个受控的生成任务:给定一个维基百科表格和一组突出显示的表格单元格,生成一个句子描述。为了获得生成的目标是自然的,但也忠实于源表,我们引入了一个数据集的建设过程中,注释者直接修改现有的候选句子从维基百科。我们对我们的数据集和注释过程进行了系统的分析,以及通过几个最先进的基线实现的结果。虽然通常很流畅,但现有的方法经常会出现表格不支持的短语,这表明该数据集可以作为高精度条件文本生成的有用研究基准。
We present ToTTo, an open-domain English table-to-text dataset with over 120,000 training examples that proposes a controlled generation task: given a Wikipedia table and a set of highlighted table cells, produce a one-sentence description. To obtain generated targets that are natural but also faithful to the source table, we introduce a dataset construction process where annotators directly revise existing candidate sentences from Wikipedia. We present systematic analyses of our dataset and annotation process as well as results achieved by several state-of-the-art baselines. While usually fluent, existing methods often hallucinate phrases that are not supported by the table, suggesting that this dataset can serve as a useful research benchmark for high-precision conditional text generation.