What’s in a Name? Answer Equivalence For Open-Domain Question Answering

What’s in a Name? Answer Equivalence For Open-Domain Question Answering
复制标题

DOI:
10.18653/v1/2021.emnlp-main.757
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Chenglei Si;Chen Zhao;Jordan L. Boyd-Graber
Chenglei Si;Chen Zhao;Jordan L. Boyd-Graber
中科院分区:
其他
文献类型:
--
作者:
Chenglei Si;Chen Zhao;Jordan L. Boyd-Graber

文献摘要

相似文献

QA评估中的一个缺陷是注释通常只提供一个金色的答案。因此,模型预测在语义上与答案相同,但表面上不同,被认为是不正确的。这项工作探索了从知识库中挖掘别名实体并将其用作额外的黄金答案(即等价答案)。我们结合了两种设置的答案:带有附加答案的评估和具有等价答案的模型培训。我们分析了三个QA基准:自然问题、TriviaQA和团队。答案扩展提高了用于评估的所有数据集的精确匹配分数,同时结合它有助于对真实数据集进行建模训练。我们通过人工事后评估确保其他答案是有效的。
A flaw in QA evaluation is that annotations often only provide one gold answer. Thus, model predictions semantically equivalent to the answer but superficially different are considered incorrect. This work explores mining alias entities from knowledge bases and using them as additional gold answers (i.e., equivalent answers). We incorporate answers for two settings: evaluation with additional answers and model training with equivalent answers. We analyse three QA benchmarks: Natural Questions, TriviaQA, and SQuAD. Answer expansion increases the exact match score on all datasets for evaluation, while incorporating it helps model training over real-world datasets. We ensure the additional answers are valid through a human post hoc evaluation.