On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert Ratings

On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert Ratings
复制标题

DOI:
10.18653/v1/2021.emnlp-main.596
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Peter Alexander Jansen;Kelly Smith;Dan Moreno;Huitzilin Ortiz
Peter Alexander Jansen;Kelly Smith;Dan Moreno;Huitzilin Ortiz
中科院分区:
其他
文献类型:
--
作者:
Peter Alexander Jansen;Kelly Smith;Dan Moreno;Huitzilin Ortiz

文献摘要

被引文献

相似文献

构建组合解释需要模型将两个或两个以上的事实联合收割机组合在一起,共同描述为什么问题的答案是正确的。通常,这些“多跳”解释相对于一个(或少量)黄金解释进行评估。在这项工作中,我们发现这些评估大大低估了模型的性能,无论是在所包含的事实的相关性,以及模型生成的解释的完整性,因为模型经常发现和产生有效的解释,是不同于黄金的解释。为了解决这个问题,我们构建了一个大型语料库的126k领域专家(科学教师)的相关性评级,增加了语料库的解释标准化的科学考试问题,发现80k额外的相关事实不评为黄金。我们基于不同的方法论(生成、排名和模式)构建了三个强大的模型,并实证表明,虽然专家增强的评级可以更好地估计解释质量,但原始(黄金)和专家增强的自动评估仍然大大低估了性能与完全手动相比,专家判断高达36%,不同的模型受到不成比例的影响。这对准确评估组合推理模型产生的解释提出了重大的方法论挑战。
Building compositional explanations requires models to combine two or more facts that, together, describe why the answer to a question is correct. Typically, these “multi-hop” explanations are evaluated relative to one (or a small number of) gold explanations. In this work, we show these evaluations substantially underestimate model performance, both in terms of the relevance of included facts, as well as the completeness of model-generated explanations, because models regularly discover and produce valid explanations that are different than gold explanations. To address this, we construct a large corpus of 126k domain-expert (science teacher) relevance ratings that augment a corpus of explanations to standardized science exam questions, discovering 80k additional relevant facts not rated as gold. We build three strong models based on different methodologies (generation, ranking, and schemas), and empirically show that while expert-augmented ratings provide better estimates of explanation quality, both original (gold) and expert-augmented automatic evaluations still substantially underestimate performance by up to 36% when compared with full manual expert judgements, with different models being disproportionately affected. This poses a significant methodological challenge to accurately evaluating explanations produced by compositional reasoning models.