A Bayesian View of Language Evolution by Iterated Learning - eScholarship

A Bayesian View of Language Evolution by Iterated Learning - eScholarship
复制标题

通过迭代学习的贝叶斯语言进化观 - eScholarship

DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
M. Kalish
M. Kalish
中科院分区:
--
文献类型:
--
作者:
T. Griffiths;M. Kalish

文献摘要

被引文献

相似文献

语言进化的贝叶斯观点迭代学习托马斯·L·格里菲斯(Tom Griffiths@Brown.edu)普罗维登斯布朗大学认知和语言科学系,RI 02912迈克尔·L·卡利什(Kalish@Louisiana.edu)路易斯安那大学拉斐特分校认知科学研究所,洛杉矶拉斐特70504(B)x 0 y0 on假设x 1 h 1 y 1数据仓库在g(A)中的数据形成所有逻辑上可能的交流方案的子集,具有所有语言共享的普遍性质(Comrie,1981);格林伯格,1963;霍金斯,1988年)。对这些语言共性的传统解释是,它们是先天的、特定于语言的、遗传的天赋对可学习语言集合施加限制的结果(例如,Chomsky,1965)。最近的研究探索了一种另类母语的解释:共性来自语言代际传递所产生的进化过程(例如,Kirby,2001;Nowak,Plotkin,&Jansen,2000)。语言随着每一代人的学习而变化。这种迭代学习的过程隐含地选择了更容易学习的语言。这提出了一个诱人的假设:迭代学习可能足以解释语言共性的出现(Briscoe,2002)。Kirby(2001)提出了一个探索这一假说的框架,称为迭代学习模型(ILM)。在ILM中,每一代人都由一个或多个学习者组成。每个学习者看到一些数据,形成关于产生这些数据的过程的假设,然后产生数据,这些数据将被提供给下一代学习者,如图1(A)所示。成功地跨代传播的语言是那些通过迭代学习造成的“信息瓶颈”的语言。如果语言的特定属性使其更容易通过这一瓶颈,那么许多代的迭代学习可能会使这些属性变得普遍。语言学习模型可以用来探索不同的语言学习观念是如何影响语言进化的。使用ILM已经检查了各种学习算法,包括启发式语法诱导器(Kirby,2001)、联想网络(Smith,Lar n in g)和语言进化模型展示了人类语言的各个方面,如构成性,如何在相互作用的主体群体中出现。本文分析了语言是如何作为一种特定形式的交互的结果发生变化的:代理人相互学习。我们证明,当学习者是有理贝叶斯代理时,这个迭代学习过程收敛到这些学习者所假设的语言的先验分布。收敛速度由每一代人所看到的数据所传达的信息量来确定;数据信息量越少,过程收敛到前一代的速度就越快。关于Kirby,&Bright ton,2003)和最小描述长度(Bright ton,2002)。使用这些算法的迭代学习产生的语言拥有人类语言最引人注目的属性之一:竞争性。在构词语言中,话语的意义是其各部分意义的函数。对这些结果的直观解释是,作文语言的规则结构意味着它们可以从更少的数据中学习,因此更有可能通过信息瓶颈。这些从信息技术学习中产生的组合性实例提出了一个重要的问题:什么语言将在许多代的迭代学习中幸存下来?虽然已经研究了使用特定学习算法的迭代学习将出现组合性的情况(Bright ton,2002;Smith等人,2003),但没有关于语言的任意属性或广泛的学习算法类别的通用结果。在本文中,我们分析了学习者是理性贝叶斯代理的情况下的ITER学习。各种学习算法可以用贝叶斯推理来表示,而贝叶斯方法是计算语言学中许多方法的基础(Manning&Sch?utze,1999)。假设学习者是贝叶斯代理,可以得出指示迭代学习将偏爱哪些语言的分析结果。特别是,我们证明了令人惊讶的结果,即迭代贝叶斯学习产生的语言上的概率分布收敛于学习者假设的先验概率分布。这意味着,一种语言被使用的渐近概率完全不取决于语言的性质,而完全由学习者的假设决定。摘要x2h2y2图1:(A)迭代学习。(B)迭代迭代贝叶斯学习中变量间的依赖关系。
A Bayesian View of Language Evolution by Iterated Learning Thomas L. Griffiths (tom griffiths@brown.edu) Department of Cognitive and Linguistic Sciences, Brown University, Providence, RI 02912 Michael L. Kalish (kalish@louisiana.edu) Institute of Cognitive Science, University of Louisiana at Lafayette, Lafayette, LA 70504 (b) x 0 y 0 on hypothesis x 1 h 1 y 1 data le ar n in g data ge n er at i hypothesis le ar n in g (a) data ge n er at i Human languages form a subset of all logically pos- sible communication schemes, with universal properties shared by all languages (Comrie, 1981; Greenberg, 1963; Hawkins, 1988). A traditional explanation for these lin- guistic universals is that they are the consequence of constraints on the set of learnable languages imposed by an innate, language-specific, genetic endowment (e.g., Chomsky, 1965). Recent research has explored an alter- native explanation: that universals emerge from evolu- tionary processes produced by the transmission of lan- guages across generations (e.g., Kirby, 2001; Nowak, Plotkin, & Jansen, 2000). Languages change as each gen- eration learns from that which preceded it. This process of iterated learning implicitly selects for languages that are more learnable. This suggests a tantalizing hypoth- esis: that iterated learning might be sufficient to explain the emergence of linguistic universals (Briscoe, 2002). Kirby (2001) introduced a framework for exploring this hypothesis, called the iterated learning model (ILM). In the ILM, each generation consists of one or more learners. Each learner sees some data, forms a hypothe- sis about the process that produced that data, and then produces the data which will be supplied to the next generation of learners, as shown in Figure 1 (a). The languages that succeed in being transmitted across gen- erations are those that pass through the “information bottleneck” imposed by iterated learning. If particular properties of languages make it easier to pass through that bottleneck, then many generations of iterated learn- ing might allow those properties to become universal. The ILM can be used to explore how different as- sumptions about language learning influence language evolution. A variety of learning algorithms have been examined using the ILM, including a heuristic gram- mar inducer (Kirby, 2001), associative networks (Smith, le ar n in g Models of language evolution have demonstrated how aspects of human language, such as compositionality, can arise in populations of interacting agents. This pa- per analyzes how languages change as the result of a particular form of interaction: agents learning from one another. We show that, when the learners are rational Bayesian agents, this process of iterated learning con- verges to the prior distribution over languages assumed by those learners. The rate of convergence is set by the amount of information conveyed by the data seen by each generation; the less informative the data, the faster the process converges to the prior. on Kirby, & Brighton, 2003), and minimum description length (Brighton, 2002). Iterated learning with these algorithms produces languages that possess one of the most compelling properties of human languages: compo- sitionality. In a compositional language, the meaning of an utterance is a function of the meaning of its parts. The intuitive explanation for these results is that the regular structure of compositional languages means that they can be learned from less data, and are thus more likely to pass through the information bottleneck. These instances of compositionality emerging from it- erated learning raise an important question: what lan- guages will survive many generations of iterated learn- ing? While the circumstances under which composi- tionality will emerge from iterated learning with specific learning algorithms have been investigated (Brighton, 2002; Smith et al., 2003), there are no general results for arbitrary properties of languages or broad classes of learning algorithms. In this paper, we analyze iter- ated learning for the case where the learners are rational Bayesian agents. A variety of learning algorithms can be formulated in terms of Bayesian inference, and Bayesian methods underlie many approaches in computational lin- guistics (Manning & Sch¨ utze, 1999). The assumption that the learners are Bayesian agents makes it possible to derive analytic results indicating which languages will be favored by iterated learning. In particular, we prove the surprising result that the probability distribution over languages resulting from iterated Bayesian learning con- verges to the prior probability distribution assumed by the learners. This implies that the asymptotic probabil- ity that a language is used does not depend at all upon the properties of the language, being determined entirely by the assumptions of the learner. Abstract x 2 h 2 y 2 Figure 1: (a) Iterated learning. (b) Dependencies among variables in iterated iterated Bayesian learning.