A Bayesian View of Language Evolution by Iterated Learning - eScholarship
A Bayesian View of Language Evolution by Iterated Learning - eScholarship
复制标题
通过迭代学习的贝叶斯语言进化观 - eScholarship
DOI:
--
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
M. Kalish
中科院分区:
文献类型:
--
作者:
T. Griffiths;M. Kalish
A Bayesian View of Language Evolution by Iterated Learning Thomas L. Griffiths (tom griffiths@brown.edu) Department of Cognitive and Linguistic Sciences, Brown University, Providence, RI 02912 Michael L. Kalish (kalish@louisiana.edu) Institute of Cognitive Science, University of Louisiana at Lafayette, Lafayette, LA 70504 (b) x 0 y 0 on hypothesis x 1 h 1 y 1 data le ar n in g data ge n er at i hypothesis le ar n in g (a) data ge n er at i Human languages form a subset of all logically pos- sible communication schemes, with universal properties shared by all languages (Comrie, 1981; Greenberg, 1963; Hawkins, 1988). A traditional explanation for these lin- guistic universals is that they are the consequence of constraints on the set of learnable languages imposed by an innate, language-specific, genetic endowment (e.g., Chomsky, 1965). Recent research has explored an alter- native explanation: that universals emerge from evolu- tionary processes produced by the transmission of lan- guages across generations (e.g., Kirby, 2001; Nowak, Plotkin, & Jansen, 2000). Languages change as each gen- eration learns from that which preceded it. This process of iterated learning implicitly selects for languages that are more learnable. This suggests a tantalizing hypoth- esis: that iterated learning might be sufficient to explain the emergence of linguistic universals (Briscoe, 2002). Kirby (2001) introduced a framework for exploring this hypothesis, called the iterated learning model (ILM). In the ILM, each generation consists of one or more learners. Each learner sees some data, forms a hypothe- sis about the process that produced that data, and then produces the data which will be supplied to the next generation of learners, as shown in Figure 1 (a). The languages that succeed in being transmitted across gen- erations are those that pass through the “information bottleneck” imposed by iterated learning. If particular properties of languages make it easier to pass through that bottleneck, then many generations of iterated learn- ing might allow those properties to become universal. The ILM can be used to explore how different as- sumptions about language learning influence language evolution. A variety of learning algorithms have been examined using the ILM, including a heuristic gram- mar inducer (Kirby, 2001), associative networks (Smith, le ar n in g Models of language evolution have demonstrated how aspects of human language, such as compositionality, can arise in populations of interacting agents. This pa- per analyzes how languages change as the result of a particular form of interaction: agents learning from one another. We show that, when the learners are rational Bayesian agents, this process of iterated learning con- verges to the prior distribution over languages assumed by those learners. The rate of convergence is set by the amount of information conveyed by the data seen by each generation; the less informative the data, the faster the process converges to the prior. on Kirby, & Brighton, 2003), and minimum description length (Brighton, 2002). Iterated learning with these algorithms produces languages that possess one of the most compelling properties of human languages: compo- sitionality. In a compositional language, the meaning of an utterance is a function of the meaning of its parts. The intuitive explanation for these results is that the regular structure of compositional languages means that they can be learned from less data, and are thus more likely to pass through the information bottleneck. These instances of compositionality emerging from it- erated learning raise an important question: what lan- guages will survive many generations of iterated learn- ing? While the circumstances under which composi- tionality will emerge from iterated learning with specific learning algorithms have been investigated (Brighton, 2002; Smith et al., 2003), there are no general results for arbitrary properties of languages or broad classes of learning algorithms. In this paper, we analyze iter- ated learning for the case where the learners are rational Bayesian agents. A variety of learning algorithms can be formulated in terms of Bayesian inference, and Bayesian methods underlie many approaches in computational lin- guistics (Manning & Sch¨ utze, 1999). The assumption that the learners are Bayesian agents makes it possible to derive analytic results indicating which languages will be favored by iterated learning. In particular, we prove the surprising result that the probability distribution over languages resulting from iterated Bayesian learning con- verges to the prior probability distribution assumed by the learners. This implies that the asymptotic probabil- ity that a language is used does not depend at all upon the properties of the language, being determined entirely by the assumptions of the learner. Abstract x 2 h 2 y 2 Figure 1: (a) Iterated learning. (b) Dependencies among variables in iterated iterated Bayesian learning.