A Study of Automatic Metrics for the Evaluation of Natural Language Explanations

A Study of Automatic Metrics for the Evaluation of Natural Language Explanations
复制标题

DOI:
10.18653/v1/2021.eacl-main.202
复制
发表时间:
2021-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Miruna Clinciu;Arash Eshghi;H. Hastie
Miruna Clinciu;Arash Eshghi;H. Hastie
中科院分区:
其他
文献类型:
--
作者:
Miruna Clinciu;Arash Eshghi;H. Hastie

文献摘要

被引文献

相似文献

随着透明度成为机器人和人工智能的关键,有必要评估提供透明度的方法,包括自动生成的自然语言(NL)解释。在这里,我们探讨了这种解释的产生与被广泛研究的自然语言生成(NLG)评估领域之间的相似之处。具体而言,我们研究了哪些NLG评估措施可以很好地解释。我们提出了ExBAN语料库:一个用于贝叶斯网络的NL解释的众包语料库。我们将人类主观评分与NLG自动测量结果进行了对比。我们发现基于嵌入的自动NLG评价方法,如BERTScore和BLEURT,与单词重叠度量(如BLEU和ROUGE)相比,与人类评分具有更高的相关性。这项工作对可解释的人工智能和透明的机器人和自主系统有影响。
As transparency becomes key for robotics and AI, it will be necessary to evaluate the methods through which transparency is provided, including automatically generated natural language (NL) explanations. Here, we explore parallels between the generation of such explanations and the much-studied field of evaluation of Natural Language Generation (NLG). Specifically, we investigate which of the NLG evaluation measures map well to explanations. We present the ExBAN corpus: a crowd-sourced corpus of NL explanations for Bayesian Networks. We run correlations comparing human subjective ratings with NLG automatic measures. We find that embedding-based automatic NLG evaluation methods, such as BERTScore and BLEURT, have a higher correlation with human ratings, compared to word-overlap metrics, such as BLEU and ROUGE. This work has implications for Explainable AI and transparent robotic and autonomous systems.