LENS: A Learnable Evaluation Metric for Text Simplification

LENS: A Learnable Evaluation Metric for Text Simplification
复制标题

DOI:
10.48550/arxiv.2212.09739
复制
发表时间:
2022-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Mounica Maddela;Yao Dou;David Heineman;Wei Xu
Mounica Maddela;Yao Dou;David Heineman;Wei Xu
中科院分区:
其他
文献类型:
--
作者:
Mounica Maddela;Yao Dou;David Heineman;Wei Xu

文献摘要

被引文献

相似文献

使用现代语言模型训练可学习度量近来成为一种很有前途的机器翻译自动评估方法。然而,现有的用于文本简化的人类评估数据集具有基于单一或过时模型的有限注释,使其不适合该方法。为了解决这些问题,我们引入了simpeval_语库,其中包含:SimpEval_past,包含24个过去系统的2.4K简化的12K人类评级;SimpEval_2022,一个具有挑战性的简化基准,包含360个简化的超过1K人类评级,包括GPT-3.5生成的文本。在SimpEval的训练下,我们提出了LENS,一个可学习的文本简化评价指标。广泛的实证结果表明,LENS与人类判断的相关性比现有的指标要好得多,为文本简化评估的未来进展铺平了道路。我们还介绍了Rank & Rate,这是一个人工评估框架,它使用交互式界面以列表方式对几个模型的简化程度进行评估,这确保了评估过程中的一致性和准确性,并用于创建SimpEval数据集。
Training learnable metrics using modern language models has recently emerged as a promising method for the automatic evaluation of machine translation. However, existing human evaluation datasets for text simplification have limited annotations that are based on unitary or outdated models, making them unsuitable for this approach. To address these issues, we introduce the SimpEval corpus that contains: SimpEval_past, comprising 12K human ratings on 2.4K simplifications of 24 past systems, and SimpEval_2022, a challenging simplification benchmark consisting of over 1K human ratings of 360 simplifications including GPT-3.5 generated text. Training on SimpEval, we present LENS, a Learnable Evaluation Metric for Text Simplification. Extensive empirical results show that LENS correlates much better with human judgment than existing metrics, paving the way for future progress in the evaluation of text simplification. We also introduce Rank & Rate, a human evaluation framework that rates simplifications from several models in a list-wise manner using an interactive interface, which ensures both consistency and accuracy in the evaluation process and is used to create the SimpEval datasets.