Internal Simulation of Behavior Has an Adaptive Advantage

Internal Simulation of Behavior Has an Adaptive Advantage
复制标题

行为的内部模拟具有适应性优势

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
J. Broekens
J. Broekens
中科院分区:
--
文献类型:
--
作者:
J. Broekens

文献摘要

被引文献

相似文献

行为的内部模拟具有适应性优势 约斯特·布罗肯斯(broekens@liacs.nl) 莱顿大学高级计算机科学研究所,荷兰莱顿,尼尔斯·玻尔韦格1号,2333 CA 模拟对自适应智能体学习的影响。我们的实验基于通过模拟智能体未来的交互来影响其动作选择。我们使用一个计算强化学习(RL)模型。我们的模型允许在其反应式操作模式的同时进行少量的预期模拟。我们的智能体生活在一个网格世界中,它必须自主学习觅食。智能体以该计算模型为大脑。我们比较不同的模拟策略在多大程度上导致不同的学习性能。 本文结构如下。我们简要讨论理论背景、我们的计算模型、实验方法和智能体系统。最后我们讨论结果并得出结论。 **摘要** 在本文中,我们检验行为的内部模拟具有强大的适应性(学习)优势这一假设。从进化的角度来看,可以合理地假设与不模拟行为的智能体相比,模拟行为的智能体具有额外的生存价值。我们展示了一个学习和决策计算模型的实验结果。我们的实验基于通过模拟智能体未来的交互来影响其动作选择。利用我们的模型,我们表明模拟对学习的这种影响导致了显著的学习优势。因为个体适应性的提高是一种进化上有利的特征,所以这对于模拟假设的进化合理性是一个相关的结果。 **关键词**:动作选择;自适应智能体;强化学习;模拟假设;计算模型 **理论背景** 交互主义(比克哈德1998年;2001年)除强化学习(萨顿和巴托,1998年)和模拟假设外,是我们方法的一个关键要素。交互主义将推理解释为智能体与其环境之间持续交互的结果。重要的是,交互可以是主动的和有准备的。主动交互准备了一组接下来可能的交互,称为交互潜能。当与环境的交互匹配时,这些潜能就会被激活,从而准备进一步的交互。主动和有准备的交互概念与模拟动作和感知的概念以及交互的联想链是兼容的,并在我们的计算模型中使用。 交互主义和模拟假设还有其他几个重要的共同假设: 1. 思维不一定是符号性的或理想的逻辑性的,这是早期认知模型的两个重要局限。动物不一定进行符号性思维(例如,见安德森,2003年;比克哈德,2001年),并且经常犯错(科恩和布卢姆,2002年;达马西奥,1994年)。 2. 感知和动作是同一枚硬币的两面,并且通过(至少)感觉运动控制区域高度相关。这避免了认知的输入 - 功能 - 输出范式,并且是与我们模型的神经元版本未来发展相关的围绕框架问题的一个重要点,但在本文中我们不详细关注这一点。 3. 这些假设与达马西奥(1994年)的思维概念密切相关,即思维是一个“仿佛”循环,涉及通过其躯体标记、情感影响评估器进行评估的模拟动作。躯体标记通过学习附着于情景的结果。三个系统至关重要地参与其中,即前额叶皮层(PFC)、躯体感觉皮层(SSC)和身体。这些标记背后的两种机制是身体回路和 **引言** 理解自适应智能体(包括自然的和人工的)中推理的本质是很重要的。例如,这种理解对于有效解决与框架问题相关的问题以及理解知识的泛化与专门化机制是必要的。推理本身假定有可推理的事物,即知识。因此,自适应智能体背景下的推理意味着(至少)有两个平行且互补的过程:首先是知识获取,即学习,其次是基于对所获取知识进行推理而产生的动作,即行为。推理本质上是做出一个明智的选择,一个由所获取知识提供信息并由智能体身体的约束所促成的选择。 模拟假设(赫斯洛,2002年)指出,思维由与环境交互的内部模拟组成。这个假设基于三个主要假设:动作模拟(动作可以被准备但不一定导致执行),感知模拟(感知可以由大脑自身产生而不一定需要外部刺激),以及预期(在真实动作与感知之间以及模拟动作与感知之间存在联想机制)。这些联想机制的持续激活构建了模拟预期交互的链条,称为隐蔽行为。这些交互影响实际行为,称为外显行为。赫斯洛(2002年)将思维等同于有意识的思考。我们使用模拟假设的一个更宽泛的定义,即:如上述所定义的推理,通过与环境交互的内部模拟而得到促进。我们对推理的定义意味着至少两个相关过程,即学习和行为。我们已经研究了
Internal Simulation of Behavior has an Adaptive Advantage Joost Broekens (broekens@liacs.nl) Leiden Institute of Advanced Computer Science, Leiden University, Niels Bohrweg 1, 2333 CA, Leiden, The Netherlands influence of simulation on the learning of an adaptive agent. Our experiments are based on biasing the agent’s action- selection by a simulation of its future interactions. We use a computational reinforcement learning (RL) model. Our model allows a small amount of anticipatory simulation concurrent with its reactive mode of operation. Our agent lives in a gridworld in which it must autonomously learn to forage. The agent has the computational model as brain . We compare to what extent different simulation strategies result in a different learning performance. This paper is structured as follows. We briefly discuss the theoretical background, our computational model, experimentation method, and agent system. We end with a discussion of our results and a conclusion. Abstract In this paper we test the hypothesis that internal simulation of behavior has a robust adaptive (learning) advantage. From an evolutionary perspective, it is plausible to assume that agents that simulate behavior have an additional survival value compared to those that do not. We present experimental results with a computational model of learning and decision- making. Our experiments are based on biasing the agent’s action-selection by a simulation of its future interactions. Using our model, we show that this influence of simulation on learning results in a significant learning advantage. Because increased individual adaptation is an evolutionary advantageous feature, this is a relevant result for the evolutionary plausibility of the simulation hypothesis. Keywords: action-selection; adaptive agent; reinforcement learning; simulation hypothesis; computational model. Theoretical Background Interactivism (Bickhard 1998; 2001) is a crucial element to our approach, besides reinforcement learning (Sutton and Barto, 1998) and the simulation hypothesis. Interactivism explains reasoning as resulting from the continuous interaction between an agent and its environment. Importantly, interactions can be active and prepared. Active interactions prepare a set of next possible interactions, referred to as interaction potentialities. These potentialities become active when interaction with the environment matches, and thus prepare further interactions. The concept of active and prepared interactions is compatible with the concept of simulated action and perception and the associative chaining of interactions, and is used in our computational model. Interactivism and the simulation hypothesis have several other important assumptions in common: 1). Thinking is not necessarily symbolic or ideally logical, two of the important limitations of earlier models of cognition. Animals do not necessarily think symbolically (see, e.g., Anderson, 2003; Bickhard, 2001), and frequently make mistakes (Cohen and Blum, 2002; Damasio, 1994). 2). Perception and action are two sides of the same coin, and highly related through (at least) sensory-motor control areas. This avoids the input-function-output paradigm of cognition and is an important point£highly related to issues surrounding the frame-problem£for future development of a neuronal version of our model. However, in this paper we do not focus on this point in detail. 3). These hypotheses closely relate to Damasio’s (1994) concept of thinking as an as if loop, involving simulated actions that are evaluated by their somatic markers, emotional impact estimators. Somatic markers are attached to outcomes of scenarios through learning. Three systems are critically involved, the prefrontal cortex (PFC), the somato-sensory cortex (SSC) and the body. The two mechanisms behind these markers are the body-loop and the Introduction It is important to understand the nature of reasoning in adaptive agents, both natural and artificial. This understanding is needed, for example, to efficiently solve problems related to the frame problem and to understand mechanisms of generalization versus specialization of knowledge. Reasoning itself assumes there is something to reason about, i.e., knowledge. Reasoning in the context of adaptive agents thus implies there are (at least) two parallel and complementary processes: first, knowledge acquisition, i.e., learning, and second the actions resulting from reasoning upon this acquired knowledge, i.e., behavior. Reasoning essentially is about making an informed choice, a choice informed by the acquired knowledge and made possible by the constraints of the body of the agent. The simulation hypothesis (Hesslow, 2002) states that thinking consists of internal simulation of interaction with the environment. This hypothesis is based upon three main assumptions: simulation of actions£actions can be prepared not necessarily resulting in execution£, simulation of perception£perceptions can be generated by the brain itself and not necessarily need external stimulation£, and anticipation£the existence of associative mechanisms both between real actions and perceptions and between simulated actions and perceptions. Continuous activation of these associative mechanisms constructs chains of simulated prospective interactions, known as covert behavior. These interactions bias actual behavior, known as overt behavior. Hesslow (2002) equates thinking with conscious thought. We use a broader definition of the simulation hypothesis, namely: reasoning, as defined above, is facilitated by the internal simulation of interaction with the environment. Our definition of reasoning implies at least two relevant processes, i.e., learning and behavior. We have studied the