What Question Answering can Learn from Trivia Nerds

What Question Answering can Learn from Trivia Nerds
复制标题

DOI:
10.18653/v1/2020.acl-main.662
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Jordan L. Boyd-Graber
Jordan L. Boyd-Graber
中科院分区:
其他
文献类型:
--
作者:
Jordan L. Boyd-Graber

文献摘要

被引文献

相似文献

除了机器回答问题的传统任务外,问答(QA)研究还创造了有趣的、具有挑战性的问题,帮助系统如何回答问题并揭示最佳系统。我们认为,创建QA数据集--以及随之而来的无处不在的排行榜--非常类似于举办一场琐事锦标赛:你写问题,让代理(无论是人类还是机器)回答问题,并宣布获胜者。然而,研究界忽视了几十年来从琐事社区创造的充满活力、公平和有效的问答比赛中辛苦学到的教训。在详细描述了现有QA数据集的问题之后,我们概述了可以转移到QA研究的关键经验-消除歧义、辨别技能和裁决争议-以及它们可能如何实施。
In addition to the traditional task of machines answering questions, question answering (QA) research creates interesting, challenging questions that help systems how to answer questions and reveal the best systems. We argue that creating a QA dataset—and the ubiquitous leaderboard that goes with it—closely resembles running a trivia tournament: you write questions, have agents (either humans or machines) answer the questions, and declare a winner. However, the research community has ignored the hard-learned lessons from decades of the trivia community creating vibrant, fair, and effective question answering competitions. After detailing problems with existing QA datasets, we outline the key lessons—removing ambiguity, discriminating skill, and adjudicating disputes—that can transfer to QA research and how they might be implemented.