A Computational Model of Word Learning from Multimodal Sensory Input

A Computational Model of Word Learning from Multimodal Sensory Input
复制标题

多模态感官输入的单词学习计算模型

DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
D. Roy
D. Roy
中科院分区:
--
文献类型:
--
作者:
D. Roy

文献摘要

被引文献

相似文献

婴儿如何分割连续的语音流来发现他们语言中的单词?目前的理论强调声学证据在发现单词边界方面的作用(Cutler 1991;Brent 1999;de Marcken 1996;Friederici&Wessels 1993;另见Bolinger&Gertsman 1957)。为了验证另一种假设,我们记录了照顾者与他们的语言前婴儿围绕共同物体进行游戏时的自然婴儿言语。我们还通过捕捉这些物体的图像来记录语音发生的视觉语境。我们使用了两个计算模型来分析数据,其中一个只处理声音记录,另一个模型整合了声音和视觉输入。这些模型是使用标准的语音和视觉处理技术实现的,使模型能够处理感觉数据。我们表明,与单独使用声学证据相比,使用视觉环境和口头输入相结合可以显著提高学习效果。这些结果证明了跨模式学习的力量,并表明婴儿可能会使用来自视觉和其他非声学背景的证据来帮助语音分割和口语单词发现。
How do infants segment continuous streams of speech to discover words of their language? Current theories emphasize the role of acoustic evidence in discovering word boundaries (Cutler 1991; Brent 1999; de Marcken 1996; Friederici & Wessels 1993; see also Bolinger & Gertsman 1957). To test an alternate hypothesis, we recorded natural infant-directed speech from caregivers engaged in play with their pre-linguistic infants centered around common objects. We also recorded the visual context in which the speech occurred by capturing images of these objects. We analyzed the data using two computational models, one of which processed only acoustic recordings, and a second model which integrated acoustic and visual input. The models were implemented using standard speech and vision processing techniques enabling the models to process sensory data. We show that using visual context in conjunction with spoken input dramatically improves learning when compared with using acoustic evidence alone. These results demonstrate the power of inter-modal learning and suggest that infants may use evidence from visual and other non-acoustic context to aid in speech segmentation and spoken word discovery.