Leveraging Large Language Models for Multiple Choice Question Answering

Leveraging Large Language Models for Multiple Choice Question Answering
复制标题

DOI:
10.48550/arxiv.2210.12353
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Joshua Robinson;Christopher Rytting;D. Wingate
Joshua Robinson;Christopher Rytting;D. Wingate
中科院分区:
其他
文献类型:
--
作者:
Joshua Robinson;Christopher Rytting;D. Wingate

文献摘要

被引文献

相似文献

虽然像GPT-3这样的大型语言模型(LLM)在零次、一次和几次设置的多项选择题回答(MCQA)任务中取得了令人印象深刻的结果,但它们通常落后于MCQA最新技术水平(SOTA)。MCQA任务传统上被呈现给LLM,如完形填空任务。LLM以问题为条件(没有相关的答案选项),其选择的选项是归一化后(长度等)分配的最高概率。一种更自然的提示方法是联合向LLM呈现问题和答案选项,并让它输出符号(例如,“A”)与其选择的答案选项相关联。这种方法允许模型显式地比较答案选项,降低计算成本,并减轻标记化方案和答案选项表示对答案选择的影响。为了使自然方法有效,它所使用的LLM必须能够将答案选项与代表它们的符号相关联。LLM需要我们所说的多选择符号绑定(MCSB)能力。这种能力因型号而异。我们发现,在20个不同的数据集上,具有高MCSB能力的模型使用自然方法比传统方法表现得更好,并且在很大程度上缩小了与SOTA的差距,这表明LLM的MCQA能力以前被低估了。
While large language models (LLMs) like GPT-3 have achieved impressive results on multiple choice question answering (MCQA) tasks in the zero, one, and few-shot settings, they generally lag behind the MCQA state of the art (SOTA). MCQA tasks have traditionally been presented to LLMs like cloze tasks. An LLM is conditioned on a question (without the associated answer options) and its chosen option is the one assigned the highest probability after normalization (for length, etc.). A more natural prompting approach is to present the question and answer options to the LLM jointly and have it output the symbol (e.g.,"A") associated with its chosen answer option. This approach allows the model to explicitly compare answer options, reduces computational costs, and mitigates the effects of tokenization scheme and answer option representations on answer selection. For the natural approach to be effective, the LLM it is used with must be able to associate answer options with the symbols that represent them. The LLM needs what we term multiple choice symbol binding (MCSB) ability. This ability varies greatly by model. We show that a model with high MCSB ability performs much better with the natural approach than with the traditional approach across 20 diverse datasets and largely closes the gap with the SOTA, suggesting that the MCQA ability of LLMs has been previously underestimated.