LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization

LongEval: Guidelines for Human Evaluation of Faithfulness in Long-form Summarization
复制标题

DOI:
10.48550/arxiv.2301.13298
复制
发表时间:
2023-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Kalpesh Krishna;Erin Bransom;Bailey Kuehl;Mohit Iyyer;Pradeep Dasigi;Arman Cohan;Kyle Lo
Kalpesh Krishna;Erin Bransom;Bailey Kuehl;Mohit Iyyer;Pradeep Dasigi;Arman Cohan;Kyle Lo
中科院分区:
其他
文献类型:
--
作者:
Kalpesh Krishna;Erin Bransom;Bailey Kuehl;Mohit Iyyer;Pradeep Dasigi;Arman Cohan;Kyle Lo

文献摘要

相似文献

虽然人工评估仍然是准确判断自动生成摘要的忠实性的最佳实践,但在评估长格式摘要时,几乎没有解决方案可以解决增加的难度和工作量。通过对162篇关于长式摘要的论文的调查,我们首先阐明了围绕长式摘要的当前人类评估实践。我们发现,这些论文中有73%没有对模型生成的摘要进行任何人工评估,而其他作品在处理长文档时面临新的困难(例如,低注释者间一致性)。出于我们的调查,我们提出了LongEval,一套指导方针,为人类评价的忠实性,在长形式的摘要,解决了以下挑战:(1)我们如何才能实现高注释者之间的一致性的忠实分数?(2)我们如何在保持准确的忠实度分数的同时最大限度地减少注释者的工作量?以及(3)人类是否从摘要和源代码片段之间的自动对齐中受益?我们在不同领域(SQualITY和PubMed)的两个长格式摘要数据集的注释研究中部署了LongEval,我们发现切换到更细的判断粒度(例如,小句级别)减少了忠实度分数中的注释者间差异(例如,std-dev从18.5到6.8)。我们还表明,分数从部分注释的细粒度单位高度相关的分数从一个完整的注释工作量(0.89肯德尔的tau使用50%的判断)。我们发布了我们的人类判断,注释模板和软件作为Python库用于未来的研究。
While human evaluation remains best practice for accurately judging the faithfulness of automatically-generated summaries, few solutions exist to address the increased difficulty and workload when evaluating long-form summaries. Through a survey of 162 papers on long-form summarization, we first shed light on current human evaluation practices surrounding long-form summaries. We find that 73% of these papers do not perform any human evaluation on model-generated summaries, while other works face new difficulties that manifest when dealing with long documents (e.g., low inter-annotator agreement). Motivated by our survey, we present LongEval, a set of guidelines for human evaluation of faithfulness in long-form summaries that addresses the following challenges: (1) How can we achieve high inter-annotator agreement on faithfulness scores? (2) How can we minimize annotator workload while maintaining accurate faithfulness scores? and (3) Do humans benefit from automated alignment between summary and source snippets? We deploy LongEval in annotation studies on two long-form summarization datasets in different domains (SQuALITY and PubMed), and we find that switching to a finer granularity of judgment (e.g., clause-level) reduces inter-annotator variance in faithfulness scores (e.g., std-dev from 18.5 to 6.8). We also show that scores from a partial annotation of fine-grained units highly correlates with scores from a full annotation workload (0.89 Kendall’s tau using 50% judgements). We release our human judgments, annotation templates, and software as a Python library for future research.