Internal Simulation of Behavior Has an Adaptive Advantage
Internal Simulation of Behavior Has an Adaptive Advantage
复制标题
行为的内部模拟具有适应性优势
DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
J. Broekens
中科院分区:
文献类型:
--
作者:
J. Broekens
Internal Simulation of Behavior has an Adaptive Advantage Joost Broekens (broekens@liacs.nl) Leiden Institute of Advanced Computer Science, Leiden University, Niels Bohrweg 1, 2333 CA, Leiden, The Netherlands influence of simulation on the learning of an adaptive agent. Our experiments are based on biasing the agent’s action- selection by a simulation of its future interactions. We use a computational reinforcement learning (RL) model. Our model allows a small amount of anticipatory simulation concurrent with its reactive mode of operation. Our agent lives in a gridworld in which it must autonomously learn to forage. The agent has the computational model as brain . We compare to what extent different simulation strategies result in a different learning performance. This paper is structured as follows. We briefly discuss the theoretical background, our computational model, experimentation method, and agent system. We end with a discussion of our results and a conclusion. Abstract In this paper we test the hypothesis that internal simulation of behavior has a robust adaptive (learning) advantage. From an evolutionary perspective, it is plausible to assume that agents that simulate behavior have an additional survival value compared to those that do not. We present experimental results with a computational model of learning and decision- making. Our experiments are based on biasing the agent’s action-selection by a simulation of its future interactions. Using our model, we show that this influence of simulation on learning results in a significant learning advantage. Because increased individual adaptation is an evolutionary advantageous feature, this is a relevant result for the evolutionary plausibility of the simulation hypothesis. Keywords: action-selection; adaptive agent; reinforcement learning; simulation hypothesis; computational model. Theoretical Background Interactivism (Bickhard 1998; 2001) is a crucial element to our approach, besides reinforcement learning (Sutton and Barto, 1998) and the simulation hypothesis. Interactivism explains reasoning as resulting from the continuous interaction between an agent and its environment. Importantly, interactions can be active and prepared. Active interactions prepare a set of next possible interactions, referred to as interaction potentialities. These potentialities become active when interaction with the environment matches, and thus prepare further interactions. The concept of active and prepared interactions is compatible with the concept of simulated action and perception and the associative chaining of interactions, and is used in our computational model. Interactivism and the simulation hypothesis have several other important assumptions in common: 1). Thinking is not necessarily symbolic or ideally logical, two of the important limitations of earlier models of cognition. Animals do not necessarily think symbolically (see, e.g., Anderson, 2003; Bickhard, 2001), and frequently make mistakes (Cohen and Blum, 2002; Damasio, 1994). 2). Perception and action are two sides of the same coin, and highly related through (at least) sensory-motor control areas. This avoids the input-function-output paradigm of cognition and is an important point£highly related to issues surrounding the frame-problem£for future development of a neuronal version of our model. However, in this paper we do not focus on this point in detail. 3). These hypotheses closely relate to Damasio’s (1994) concept of thinking as an as if loop, involving simulated actions that are evaluated by their somatic markers, emotional impact estimators. Somatic markers are attached to outcomes of scenarios through learning. Three systems are critically involved, the prefrontal cortex (PFC), the somato-sensory cortex (SSC) and the body. The two mechanisms behind these markers are the body-loop and the Introduction It is important to understand the nature of reasoning in adaptive agents, both natural and artificial. This understanding is needed, for example, to efficiently solve problems related to the frame problem and to understand mechanisms of generalization versus specialization of knowledge. Reasoning itself assumes there is something to reason about, i.e., knowledge. Reasoning in the context of adaptive agents thus implies there are (at least) two parallel and complementary processes: first, knowledge acquisition, i.e., learning, and second the actions resulting from reasoning upon this acquired knowledge, i.e., behavior. Reasoning essentially is about making an informed choice, a choice informed by the acquired knowledge and made possible by the constraints of the body of the agent. The simulation hypothesis (Hesslow, 2002) states that thinking consists of internal simulation of interaction with the environment. This hypothesis is based upon three main assumptions: simulation of actions£actions can be prepared not necessarily resulting in execution£, simulation of perception£perceptions can be generated by the brain itself and not necessarily need external stimulation£, and anticipation£the existence of associative mechanisms both between real actions and perceptions and between simulated actions and perceptions. Continuous activation of these associative mechanisms constructs chains of simulated prospective interactions, known as covert behavior. These interactions bias actual behavior, known as overt behavior. Hesslow (2002) equates thinking with conscious thought. We use a broader definition of the simulation hypothesis, namely: reasoning, as defined above, is facilitated by the internal simulation of interaction with the environment. Our definition of reasoning implies at least two relevant processes, i.e., learning and behavior. We have studied the