One-shot learning of generative speech concepts

One-shot learning of generative speech concepts
复制标题

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
2.5
通讯作者:
B. Lake;Chia-ying Lee;James R. Glass;J. Tenenbaum
B. Lake;Chia-ying Lee;James R. Glass;J. Tenenbaum
中科院分区:
心理学3区
文献类型:
--
作者:
B. Lake;Chia-ying Lee;James R. Glass;J. Tenenbaum

文献摘要

被引文献

相似文献

生成语音概念的一次性学习Brenden M. Lake* Chia-ying Lee* James R. Glass Joshua B. Tenenbaum脑与认知科学MIT CSAIL MIT CSAIL MIT脑与认知科学MIT Abstract 2007)。相关的计算工作调查了其他有助于学习词义的因素,包括学习学习哪些特征是重要的(Colunga & Smith, 2005; Kemp et al., 2007)和跨情境的单词学习(Smith & Yu, 2008; Frank, Goodman, & Tenen- baum, 2009)。但无论如何,意义的习得是可能的,因为孩子们也可以把这个词作为一个类别来学习,把一个词的所有实例(不包括非实例),比如“大象”,映射到相同的语音表示,而不考虑说话人的身份和其他声学变化的来源。这是本文的重点。先前的研究表明,儿童可以进行一次口语单词学习(Carey & Bartlett, 1978)。当孩子们(3-4岁)被要求拿出一个“铬”色的物体时,他们似乎把这个声音标记为一个新词;有些人后来甚至产生了自己的“铬”一词的近义词。此外,无论是学习第二语言,一个新名字,还是一个新的词汇,学习新的口语单词在成年后仍然是一个重要的问题。我们工作的目标是双重的:开发一次性学习任务,可以并排比较人和模型,并开发一个在这些任务上表现良好的计算模型。由于这些任务对人和算法来说都必须包含新单词,我们测试了说英语的人学习日语单词的能力。这种语言配对也为通过语音结构迁移学习到学习提供了一个有趣的测试案例,因为日语与英语音素的相似之处大致属于英语音素的一个子集(Ohata, 2004)。单次学习模型的最新进展能否用于从原始语音中学习新的口语单词?如何从一个例子中学习一个单词的生成模型?最近的行为和计算研究表明,组合性与分层贝叶斯建模相结合,可以成为构建“生成模型的生成模型”的强大方法,支持一次性学习(Lake, Salakhutdinov, & Tenenbaum, 2012; Lake等人,2013)。这个想法被应用于一次性学习手写字符,这是一种类似的高维自然概念,使用“综合分析”方法。给定一个小说人物的原始图像,该模型学习通过一个潜在的动态因果过程来表示它,该过程由笔画及其空间关系组成(图1a)。跨概念共享随机马达原语(图1a-i)提供了一种从现有模型中合成新生成模型的方法(图1a-iii)。合成生成模型非常适合口语单词习得问题,因为它们与经典的一次性学习(人类从一个或几个例子中学习新概念的能力)有关,这对传统的学习算法构成了挑战,尽管基于高层次贝叶斯模型和合成表示的方法已经取得了进展。本文研究了儿童和成人如何从一个例子中轻松地学习新单词的口语形式-识别新语音序列的任意实例,并排除非实例,而不考虑说话者身份和声学变化。这是学习单词含义和学习使用它的重要一步,我们开发了一个层次贝叶斯声学模型,可以从一个例子中学习口语单词,利用无监督学习的产物音素单位的组合。我们比较了人类和计算模型在一次性日语新单词分类和生成任务上的表现,发现学习单元在取得良好的性能方面起着重要作用。关键词:一次性学习;语音识别;类别学习;人们可以从一个或几个例子中学习一个新概念,做出有意义的概括,远远超出观察到的数据。在机器中复制这种能力是具有挑战性的,因为标准的学习算法需要数十、数百或数千个示例才能达到高水平的分类性能。尽管如此,最近对认知科学和机器学习的兴趣提高了我们对“一次性学习”的计算理解,并出现了几个关键主题。概率生成模型可以仅从一个或几个例子中预测人们的一般情况,如低维空间中的数据所示(Shepard, 1987; Tenenbaum & Griffiths, 2001)。另一个主题是围绕“学习到学习”发展起来的,即一次性学习本身是由之前的相关概念学习发展而来的,而高层次贝叶斯(HB)模型可以通过高光化对泛化最重要的维度或特征来学习(Fei-Fei, Fergus, & Perona, 2006; Kemp, Perfors, & Tenenbaum, 2007; Salakhutdinov, Tenenbaum, & Torralba, 2012)。在本文中,我们研究了学习新口语的问题,这是语言发展的一个重要组成部分。据估计,孩子们从一岁到高中毕业平均每天学习10个新单词(Bloom, 2000)。为了以如此惊人的速度学习,孩子们必须从很少的数据中学习新单词。以前的计算工作主要集中在从几个例子中学习单词的意义;例如,当听到“大象”这个词与一个前例子配对时,孩子必须决定哪些物体属于“大象”组,哪些不属于(例如,徐和特南鲍姆,*前两位作者对这项工作的贡献相同。
One-shot learning of generative speech concepts Brenden M. Lake* Chia-ying Lee* James R. Glass Joshua B. Tenenbaum Brain and Cognitive Sciences MIT CSAIL MIT CSAIL MIT Brain and Cognitive Sciences MIT Abstract 2007). Related computational work has investigated other factors that contribute to learning word meaning, including learning-to-learn which features are important (Colunga & Smith, 2005; Kemp et al., 2007) and cross-situational word learning (Smith & Yu, 2008; Frank, Goodman, & Tenen- baum, 2009). But by any account, the acquisition of mean- ing is only possible because the child can also learn the spo- ken word as a category, mapping all instances (and exclud- ing non-instances) of a word like “elephant” to the same phonological representation, regardless of speaker identify and other sources of acoustic variability. This is the focus of the current paper. Previous work has shown that chil- dren can do one-shot spoken word learning (Carey & Bartlett, 1978). When children (ages 3-4) were asked to bring over a “chromium” colored object, they seemed to flag the sound as a new word; some even later produced their own approxima- tion of the word “chromium.” Furthermore, acquiring new spoken words remains an important problem well into adult- hood whether its learning a second language, a new name, or a new vocabulary word. The goal of our work is twofold: to develop one-shot learn- ing tasks that can compare people and models side-by-side, and to develop a computational model that performs well on these tasks. Since the tasks must contain novel words for both people and algorithms, we tested English speakers on their ability to learn Japanese words. This language pairing also offers an interesting test case for learning-to-learn through the transfer of phonetic structure, since the Japanese analogs to English phonemes fall roughly within a subset of English phonemes (Ohata, 2004). Can the recent progress on models of one-shot learning be leveraged for learning new spoken words from raw speech? How could a generative model of a word be learned from just one example? Recent behavioral and computational work suggests that compositionality, combined with Hierarchical Bayesian modeling, can be a powerful way to build a “gen- erative model for generative models” that supports one-shot learning (Lake, Salakhutdinov, & Tenenbaum, 2012; Lake et al., 2013). This idea was applied to the one-shot learning of handwritten characters, a similarly high-dimensional do- main of natural concepts, using an “analysis-by-synthesis” approach. Given a raw image of a novel character, the model learns to represent it by a latent dynamic causal process, com- posed of pen strokes and their spatial relations (Fig. 1a). The sharing of stochastic motor primitives across concepts (Fig. 1a-i) provides a means of synthesizing new generative mod- els out of pieces of existing ones (Fig. 1a-iii). Compositional generative models are well-suited for the problem of spoken word acquisition, as they relate to classic One-shot learning – the human ability to learn a new concept from just one or a few examples – poses a challenge to tradi- tional learning algorithms, although approaches based on Hi- erarchical Bayesian models and compositional representations have been making headway. This paper investigates how chil- dren and adults readily learn the spoken form of new words from one example – recognizing arbitrary instances of a novel phonological sequence, and excluding non-instances, regard- less of speaker identity and acoustic variability. This is an es- sential step on the way to learning a word’s meaning and learn- ing to use it, and we develop a Hierarchical Bayesian acoustic model that can learn spoken words from one example, utiliz- ing compositions of phoneme-like units that are the product of unsupervised learning. We compare people and computa- tional models on one-shot classification and generation tasks with novel Japanese words, finding that the learned units play an important role in achieving good performance. Keywords: one-shot learning; speech recognition; category learning; exemplar generation Introduction People can learn a new concept from just one or a few ex- amples, making meaningful generalizations that go far be- yond the observed data. Replicating this ability in machines has been challenging, since standard learning algorithms re- quire tens, hundreds, or thousands of examples before reach- ing a high level of classification performance. Nonetheless, recent interest from cognitive science and machine learning has advanced our computational understanding of “one-shot learning,” and several key themes have emerged. Proba- bilistic generative models can predict how people general- ize from just one or a few examples, as shown for data ly- ing in a low-dimensional space (Shepard, 1987; Tenenbaum & Griffiths, 2001). Another theme has developed around learning-to-learn, the idea that one-shot learning itself de- velops from previous learning with related concepts, and Hi- erarchical Bayesian (HB) models can learn-to-learn by high- lighting the dimensions or features that are most important for generalization (Fei-Fei, Fergus, & Perona, 2006; Kemp, Perfors, & Tenenbaum, 2007; Salakhutdinov, Tenenbaum, & Torralba, 2012). In this paper, we study the problem of learning new spoken words, an essential ingredient for language development. By one estimate, children learn an average of ten new words per day from the age of one to the end of high school (Bloom, 2000). For learning to proceed at such an astounding rate, children must be learning new words from very little data. Previous computational work has focused on the problem of learning the meaning of words from a few examples; for in- stance, upon hearing the word “elephant” paired with an ex- emplar, the child must decide which objects belong to the set of “elephants” and which do not (e.g., Xu & Tenenbaum, * The first two authors contributed equally to this work.