Constructing a hypothesis space from the Web for large-scale Bayesian word learning
Constructing a hypothesis space from the Web for large-scale Bayesian word learning
复制标题
从网络构建用于大规模贝叶斯单词学习的假设空间
DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Thomas L. Griffiths
中科院分区:
文献类型:
--
作者:
Joshua T. Abbott;Joseph L. Austerweil;Thomas L. Griffiths
Constructing a hypothesis space from the Web for large-scale Bayesian word learning Joshua T. Abbott (joshua.abbott@berkeley.edu) Joseph L. Austerweil (joseph.austerweil@gmail.com) Thomas L. Griffiths (tom griffiths@berkeley.edu) Department of Psychology, University of California, Berkeley, CA 94720 USA Abstract Bayesian generalization model. In this paper, we use this approach to show how a hypothesis space and prior can be constructed automatically from a large online database, mak- ing it possible to apply the Bayesian generalization frame- work to a wide range of naturalistic stimuli. We focus on one specific generalization problem, word learning, where peo- ple learn new words from observing a few objects that can be labeled with that word. Given that the number of possible ex- tensions of a word is essentially infinite, learning the objects referred to by a word is a very difficult inductive problem (Quine, 1975). Xu and Tenenbaum (2007) showed how the Bayesian generalization framework could be used to explain how people learn new words. However, to construct the hy- pothesis space of their Bayesian model, Xu and Tenenbaum (2007) elicited approximately 400 similarity judgments from their participants. Clearly this is not practical to extend into every domain where people learn words. Thus, word learn- ing is an appropriate setting for exploring novel methods of constructing hypothesis spaces and prior distributions. The Bayesian generalization framework has been successful in explaining how people generalize a property from a few observed stimuli to novel stimuli, across several different domains. To create a successful Bayesian generalization model, modelers typically specify a hypothesis space and prior probability distribution for each specific domain. How- ever, this raises two problems: the models do not scale beyond the (typically small-scale) domain that they were designed for, and the explanatory power of the models is reduced by their reliance on a hand-coded hypothesis space and prior. To solve these two problems, we propose a method for deriving hypothesis spaces and priors from large online databases. We evaluate our method by constructing a hypothesis space and prior for a Bayesian word learning model from WordNet, a large online database that encodes the semantic relationships between words as a network. After validating our approach by replicating a previous word learning study, we apply the same model to a new experiment featuring three additional taxonomic domains (clothing, containers, and seats). In both experiments, we found that the same automatically constructed hypothesis space explains the complex pattern of generalization behavior, producing accurate predictions across a total of six different domains. We propose a method for automatically constructing the hypothesis space and prior distribution of a Bayesian word learning model using freely available online resources. In particular, we use WordNet (Fellbaum, 2010; Miller, 1995) as an initial source for automatically creating the hypothesis space, and ImageNet (Deng et al., 2009) as a source of natu- ralistic images that can be used as stimuli to test the resulting model in behavioral experiments. WordNet is a popular lexi- cal database of English comprised of over 100,000 relational sets of synonyms. ImageNet is a large ontology of images conforming to the hierarchical structure of WordNet, with the aim of providing over 500 high-quality images per noun in WordNet. These resources allow us to construct hypothesis spaces and prior distributions for word learning without elic- iting a single judgment from participants and test the result- ing model on a much larger scale than was previously pos- sible. We demonstrate that the Bayesian model formulated from WordNet captures participant judgments in two behav- ioral experiments, addressing the practical and theoretical is- sues with Bayesian models discussed earlier. Keywords: generalization; concept learning; word learning; Bayesian modeling; online databases Introduction Many problems solved by the mind conform to the same ab- stract computational formulation: How should a property be generalized to novel stimuli from a set of stimuli observed to have the property? As there are many ways to extend the property that are consistent with some observed evidence, these are problems of induction, where the evidence con- strains, but does not determine, the solution to a problem. The Bayesian generalization framework (Shepard, 1987; Tenen- baum & Griffiths, 2001) has been remarkably successful at explaining human generalization behavior in a wide range of domains. However, its success is largely dependent on the choice of a hypothesis space and a prior probability distribu- tion on hypotheses, which are usually hand constructed by the researcher for each specific problem. This is unsatisfy- ing practically, because the models do not scale beyond the originally modeled problem, and theoretically, as it is unclear whether their success is due to the cleverness of the modeler and not because of a deep mathematical property of the com- putational problem that people solve. One possible solution is to use existing sources of infor- mation about the organization of a domain as the basis for specifying a hypothesis space and prior. This helps address both the practical and the theoretical concerns raised by the The plan of the rest of the paper is as follows. In the next sections we review the Bayesian generalization model and then examine how Xu and Tenenbaum (2007) constructed the hypothesis space for their Bayesian word learning model. We then show how to build a hypothesis space from WordNet that can be used to evaluate word learning models on a large scale. Afterwards, we present two experiments utilizing this hypoth-