Emergent analogical reasoning in large language models

Emergent analogical reasoning in large language models
复制标题

大型语言模型中的涌现类比推理

DOI:
10.1038/s41562-023-01659-w
复制
发表时间:
2022-12
影响因子:
29.9
通讯作者:
Taylor W. Webb;K. Holyoak;Hongjing Lu
Taylor W. Webb;K. Holyoak;Hongjing Lu
中科院分区:
心理学1区
文献类型:
--
作者:
Taylor W. Webb;K. Holyoak;Hongjing Lu

文献摘要

相似文献

最近大型语言模型的出现再次引发了关于人类认知能力是否可能在给定足够训练数据的通用模型中出现的争论。特别令人感兴趣的是这些模型在没有任何直接训练的情况下对新问题进行零射击推理的能力。在人类的认知中,这种能力与类比推理的能力密切相关。在这里,我们在一系列类比任务上对人类推理器和大型语言模型(生成预训练转换器(GPT)-3的文本-davinci-003变体)进行了直接比较,包括基于Raven标准渐进矩阵规则结构的非视觉矩阵推理任务。我们发现,GPT-3表现出惊人的强大的抽象模式归纳能力,在大多数情况下匹配甚至超过人类的能力;GPT-4的初步测试显示,其性能甚至更好。我们的研究结果表明,像GPT-3这样的大型语言模型已经获得了一种紧急能力,可以为广泛的类比问题找到零射击解决方案。
The recent advent of large language models has reinvigorated debate over whether human cognitive capacities might emerge in such generic models given sufficient training data. Of particular interest is the ability of these models to reason about novel problems zero-shot, without any direct training. In human cognition, this capacity is closely tied to an ability to reason by analogy. Here we performed a direct comparison between human reasoners and a large language model (the text-davinci-003 variant of Generative Pre-trained Transformer (GPT)-3) on a range of analogical tasks, including a non-visual matrix reasoning task based on the rule structure of Raven’s Standard Progressive Matrices. We found that GPT-3 displayed a surprisingly strong capacity for abstract pattern induction, matching or even surpassing human capabilities in most settings; preliminary tests of GPT-4 indicated even better performance. Our results indicate that large language models such as GPT-3 have acquired an emergent ability to find zero-shot solutions to a broad range of analogy problems.