MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers

MAUVE: Measuring the Gap Between Neural Text and Human Text using Divergence Frontiers
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
--
影响因子:
--
通讯作者:
Krishna Pillutla;Swabha Swayamdipta;Rowan Zellers;John Thickstun;S. Welleck;Yejin Choi;Zaïd Harchaoui
Krishna Pillutla;Swabha Swayamdipta;Rowan Zellers;John Thickstun;S. Welleck;Yejin Choi;Zaïd Harchaoui
中科院分区:
其他
文献类型:
--
作者:
Krishna Pillutla;Swabha Swayamdipta;Rowan Zellers;John Thickstun;S. Welleck;Yejin Choi;Zaïd Harchaoui

文献摘要

相似文献

随着开放式文本生成取得重大进展,衡量机器生成的文本与人类语言的接近程度仍然是一个关键的开放性问题。我们引入了MAUVE,这是一种用于开放式文本生成的比较度量,它使用散度边界直接将文本生成模型所学习到的分布与人类撰写的文本的分布进行比较。MAUVE通过在量化嵌入空间中计算信息散度,可扩展到现代文本生成模型。通过对三个开放式生成任务进行广泛的实证研究,我们发现MAUVE能够识别生成文本的已知属性,随模型大小自然扩展,并且与人类判断相关,相比现有的分布评估指标限制更少。
As major progress is made in open-ended text generation, measuring how close machine-generated text is to human language remains a critical open problem. We introduce MAUVE, a comparison measure for open-ended text generation, which directly compares the learnt distribution from a text generation model to the distribution of human-written text using divergence frontiers. MAUVE scales up to modern text generation models by computing information divergences in a quantized embedding space. Through an extensive empirical study on three open-ended generation tasks, we find that MAUVE identifies known properties of generated text, scales naturally with model size, and correlates with human judgments, with fewer restrictions than existing distributional evaluation metrics.