Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task

Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
复制标题

DOI:
10.48550/arxiv.2210.13382
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Kenneth Li;Aspen K. Hopkins;David Bau;Fernanda Vi'egas;H. Pfister;M. Wattenberg
Kenneth Li;Aspen K. Hopkins;David Bau;Fernanda Vi'egas;H. Pfister;M. Wattenberg
中科院分区:
其他
文献类型:
--
作者:
Kenneth Li;Aspen K. Hopkins;David Bau;Fernanda Vi'egas;H. Pfister;M. Wattenberg

文献摘要

被引文献

相似文献

语言模型展示了令人惊讶的能力范围,但他们明显能力的来源尚不清楚。这些网络只是记住了表面统计数据的集合,还是依赖于生成它们所看到的序列的过程的内部表示?我们通过将GPT模型的一个变体应用到一个简单的棋盘游戏Othello中预测合法移动的任务来研究这个问题。尽管网络没有关于游戏或其规则的先验知识,但我们发现了董事会状态的紧急非线性内部表示的证据。干预性实验表明,这种表征可以用来控制网络的输出,并创建有助于用人类术语解释预测的“潜在显著图”。
Language models show a surprising range of capabilities, but the source of their apparent competence is unclear. Do these networks just memorize a collection of surface statistics, or do they rely on internal representations of the process that generates the sequences they see? We investigate this question by applying a variant of the GPT model to the task of predicting legal moves in a simple board game, Othello. Although the network has no a priori knowledge of the game or its rules, we uncover evidence of an emergent nonlinear internal representation of the board state. Interventional experiments indicate this representation can be used to control the output of the network and create"latent saliency maps"that can help explain predictions in human terms.