Neural Foundations of Mental Simulation: Future Prediction of Latent Representations on Dynamic Scenes

Neural Foundations of Mental Simulation: Future Prediction of Latent Representations on Dynamic Scenes
复制标题

DOI:
10.48550/arxiv.2305.11772
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Aran Nayebi;R. Rajalingham;M. Jazayeri;G. R. Yang
Aran Nayebi;R. Rajalingham;M. Jazayeri;G. R. Yang
中科院分区:
其他
文献类型:
--
作者:
Aran Nayebi;R. Rajalingham;M. Jazayeri;G. R. Yang

文献摘要

相似文献

人类和动物对物理世界有着丰富而灵活的理解,这使他们能够推断物体和事件的潜在动态轨迹、可能的未来状态,并利用这些来计划和预测行动的后果。然而,这些计算背后的神经机制尚不清楚。我们将目标驱动的建模方法与密集的神经生理学数据和包含数千个比较的高通量人类行为读数相结合,以直接解决这个问题。具体来说,我们构建并评估了几类感觉认知网络,以预测丰富的、与行为学相关的环境的未来状态,范围从具有像素级或对象槽目标的自监督端到端模型,到在纯静态图像预训练或动态视频预训练基础模型的潜在空间中进行未来预测的模型。我们发现“规模并不是你所需要的”,许多最先进的机器学习模型无法在我们未来预测的神经和行为基准上表现良好。事实上,只有一类模型总体上与这些数据匹配良好。我们发现,目前神经反应最好的预测方法是通过经过训练的模型来预测其环境的未来状态,这些模型在以自我监督的方式针对动态场景进行优化的预训练基础模型的潜在空间中。这些模型还接近神经元预测视觉上隐藏的环境状态变量的能力,尽管没有经过明确的训练。最后,我们发现并非所有基础模型潜伏都是相等的。值得注意的是,未来在视频基础模型的潜在空间中进行预测的模型经过优化以支持各种以自我为中心的感觉运动任务,在我们能够测试的所有环境场景中合理地匹配人类行为错误模式和神经动力学。总体而言,这些发现表明,灵长类心理模拟的神经机制和行为具有与之相关的强烈归纳偏差,并且迄今为止与优化以预测未来对可重用视觉表示的预测最一致,这些视觉表示对更普遍的实体人工智能有用。
Humans and animals have a rich and flexible understanding of the physical world, which enables them to infer the underlying dynamical trajectories of objects and events, plausible future states, and use that to plan and anticipate the consequences of actions. However, the neural mechanisms underlying these computations are unclear. We combine a goal-driven modeling approach with dense neurophysiological data and high-throughput human behavioral readouts that contain thousands of comparisons to directly impinge on this question. Specifically, we construct and evaluate several classes of sensory-cognitive networks to predict the future state of rich, ethologically-relevant environments, ranging from self-supervised end-to-end models with pixel-wise or object-slot objectives, to models that future predict in the latent space of purely static image-pretrained or dynamic video-pretrained foundation models. We find that “scale is not all you need”, and that many state-of-the-art machine learning models fail to perform well on our neural and behavioral benchmarks for future prediction. In fact, only one class of models matches these data well overall. We find that neural responses are currently best predicted by models trained to predict the future state of their environment in the latent space of pretrained foundation models optimized for dynamic scenes in a self-supervised manner. These models also approach the neurons’ ability to predict the environmental state variables that are visually hidden from view, despite not being explicitly trained to do so. Finally, we find that not all foundation model latents are equal. Notably, models that future predict in the latent space of video foundation models that are optimized to support a diverse range of egocentric sensorimotor tasks, reasonably match both human behavioral error patterns and neural dynamics across all environmental scenarios that we were able to test. Overall, these findings suggest that the neural mechanisms and behaviors of primate mental simulation have strong inductive biases associated with them, and are thus far most consistent with being optimized to future predict on reusable visual representations that are useful for Embodied AI more generally.