Chain of Thought Prompting Elicits Reasoning in Large Language Models

Chain of Thought Prompting Elicits Reasoning in Large Language Models
复制标题

DOI:
--
复制
发表时间:
2022-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Jason Wei;Xuezhi Wang;Dale Schuurmans;Maarten Bosma;E. Chi;F. Xia;Quoc Le;Denny Zhou
Jason Wei;Xuezhi Wang;Dale Schuurmans;Maarten Bosma;E. Chi;F. Xia;Quoc Le;Denny Zhou
中科院分区:
其他
文献类型:
--
作者:
Jason Wei;Xuezhi Wang;Dale Schuurmans;Maarten Bosma;E. Chi;F. Xia;Quoc Le;Denny Zhou

文献摘要

被引文献

相似文献

我们探讨了如何生成思维链——一系列中间推理步骤——显著提高大型语言模型执行复杂推理的能力。特别是,我们通过一种称为思维链提示的简单方法,展示了这种推理能力是如何在足够大的语言模型中自然出现的,其中提供了一些思维链演示作为提示的示例。在三个大型语言模型上的实验表明,思维链提示提高了一系列算术、常识和符号推理任务的表现。经验上的收获可能是惊人的。例如,仅使用8个思维链示例提示540b参数的语言模型,在数学单词问题的GSM8K基准测试中达到了最先进的精度,甚至超过了经过微调的带有验证器的GPT-3。
We explore how generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning. In particular, we show how such reasoning abilities emerge naturally in sufficiently large language models via a simple method called chain of thought prompting, where a few chain of thought demonstrations are provided as exemplars in prompting. Experiments on three large language models show that chain of thought prompting improves performance on a range of arithmetic, commonsense, and symbolic reasoning tasks. The empirical gains can be striking. For instance, prompting a 540B-parameter language model with just eight chain of thought exemplars achieves state of the art accuracy on the GSM8K benchmark of math word problems, surpassing even finetuned GPT-3 with a verifier.