DREAM: Uncovering Mental Models behind Language Models

DREAM: Uncovering Mental Models behind Language Models
复制标题

梦想:揭示语言模型背后的心理模型

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
Peter Clark
Peter Clark
中科院分区:
--
文献类型:
--
作者:
Yuling Gu;Bhavana Dalvi;Peter Clark

文献摘要

参考文献

被引文献

相似文献

语言模型(LM)在回答情境问题时在多大程度上建立了场景的“心理模型”(例如,(一)具体的伦理问题?虽然认知科学已经表明,心智模型在人类解决问题的过程中发挥着重要作用,但目前尚不清楚现有的LMs在回答问题时的高表现是否得到了类似模型构建的支持-如果不是,这是否可以解释它们众所周知的灾难性失败。我们观察到,金刚鹦鹉,现有的T5为基础的LM,当探测提供了一些有用的,但不充分的心理模型的情景问题(估计准确率= 43%,有用性= 21%,一致性=42%)。我们提出了DREAM,一个以情景问题为输入的模型,以产生一个阐述情景的心理模型,而没有任何额外的任务特定训练数据。它通过现有NLP资源的远程监督来继承其社会常识。我们的分析表明,与金刚鹦鹉相比,梦想可以产生更好的心理模型(估计准确率= 67%,有用性= 37%,一致性=71%)。最后,由DREAM生成的心理模型可以用作情境QA任务的附加上下文。这个额外的上下文在三个不同的数据集上将Macaw零射击模型的答案准确度提高了+1%到+4%(绝对值)。
To what extent do language models (LMs) build “mental models” of a scene when answering situated questions (e.g., questions about a specific ethical dilemma)? While cognitive science has shown that mental models play a fundamental role in human problem-solving, it is unclear whether the high question-answering performance of existing LMs is backed by similar model building - and if not, whether that can explain their well-known catastrophic failures. We observed that Macaw, an existing T5-based LM, when probed provides somewhat useful but inadequate mental models for situational questions (estimated accuracy=43%, usefulness=21%, consistency=42%). We propose DREAM, a model that takes a situational question as input to produce a mental model elaborating the situation, without any additional task spe-cific training data for mental models. It inherits its social commonsense through distant supervision from existing NLP resources. Our analysis shows that DREAM can produce significantly better mental models (estimated accuracy=67%, usefulness=37%, consistency=71%) compared to Macaw. Finally, mental models generated by DREAM can be used as additional context for situational QA tasks. This additional context improves the answer accuracy of a Macaw zero-shot model by between +1% and +4% (absolute) on three different datasets.
DOI: --
发表时间: 2020-08
期刊: ArXiv
影响因子: --
作者:
Dan Hendrycks;Collin Burns;Steven Basart;Andrew Critch;J. Li;D. Song;J. Steinhardt
通讯作者: Dan Hendrycks;Collin Burns;Steven Basart;Andrew Critch;J. Li;D. Song;J. Steinhardt
DOI: 10.18653/v1/2020.emnlp-main.530
发表时间: 2020-10
期刊: --
影响因子: --
作者:
Wei-Jen Ko;Tengyang Chen;Yiyan Huang;Greg Durrett;Junyi Jessy Li
通讯作者: Wei-Jen Ko;Tengyang Chen;Yiyan Huang;Greg Durrett;Junyi Jessy Li