It’s Commonsense, isn’t it? Demystifying Human Evaluations in Commonsense-Enhanced NLG Systems

It’s Commonsense, isn’t it? Demystifying Human Evaluations in Commonsense-Enhanced NLG Systems
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Miruna Clinciu;Dimitra Gkatzia;Saad Mahamood
Miruna Clinciu;Dimitra Gkatzia;Saad Mahamood
中科院分区:
其他
文献类型:
--
作者:
Miruna Clinciu;Dimitra Gkatzia;Saad Mahamood

文献摘要

相似文献

常识是人类认知中不可或缺的一部分,它使我们能够做出正确的决定,有效地与他人沟通,并解释情况和话语。赋予人工智能系统常识知识能力将有助于我们更接近于创造展示人类智能的系统。最近在自然语言生成(NLG)方面的努力侧重于通过大规模预先训练的语言模型或通过纳入外部知识库来纳入常识知识。这样的系统在训练集中没有明确编码常识的情况下表现出推理能力。这些系统需要仔细评估,因为它们在培训期间纳入了额外的资源,这增加了更多的错误来源。此外,人类对这类系统的评估可能会有很大的差异,从而无法比较不同的系统和定义基线。本文旨在通过提出常识评估卡(CEC)来揭开常识增强型NLG系统的人类评估的神秘面纱,CEC是一套针对常识增强型NLG系统的评估报告的建议,其基础是对最近文献中报道的人类评估的广泛分析。
Common sense is an integral part of human cognition which allows us to make sound decisions, communicate effectively with others and interpret situations and utterances. Endowing AI systems with commonsense knowledge capabilities will help us get closer to creating systems that exhibit human intelligence. Recent efforts in Natural Language Generation (NLG) have focused on incorporating commonsense knowledge through large-scale pre-trained language models or by incorporating external knowledge bases. Such systems exhibit reasoning capabilities without common sense being explicitly encoded in the training set. These systems require careful evaluation, as they incorporate additional resources during training which adds additional sources of errors. Additionally, human evaluation of such systems can have significant variation, making it impossible to compare different systems and define baselines. This paper aims to demystify human evaluations of commonsense-enhanced NLG systems by proposing the Commonsense Evaluation Card (CEC), a set of recommendations for evaluation reporting of commonsense-enhanced NLG systems, underpinned by an extensive analysis of human evaluations reported in the recent literature.