Computational models of child language learning: an introduction.

Computational models of child language learning: an introduction.
复制标题

儿童语言学习的计算模型:简介。

DOI:
10.1017/s0305000910000139
复制
发表时间:
2010
影响因子:
2.2
通讯作者:
Macwhinney,Brian
Macwhinney,Brian
中科院分区:
人文科学3区
文献类型:
--
作者:
Macwhinney,Brian

文献摘要

被引文献

相似文献

这期特刊展示了儿童语言习得的计算建模方面的最新工作。总之,这九篇论文证明了语料库驱动的计算建模的科学价值。每一篇论文都提出了一个清晰的机制模型,可以根据公开的数据和/或可复制的实验进行测试,改进或拒绝。然而,对于许多读者来说,这类工作中涉及的多种形式主义和技术细节可能会成为评估所作贡献性质的障碍。因此,我的目标是在这篇介绍中总结我所看到的九个项目中每个项目所传递的重要信息。希望这个概述能鼓励读者转向九个贡献中的每一个的细节。在这里包括的九篇论文中,有八篇是基于通过儿童语言数据交换系统(CHILDES)提供的自发成人与儿童互动的语料库进行分析的。这些CHILDES语料库提供了两种对建模企业至关重要的信息。首先,他们详细记录了儿童语言的自然发展。其次,这些语料库提供了一个很好的成人语音样本,作为儿童语言学习机制的输入。考虑到这两个经验基础,计算建模者的工作是确定一组算法,这些算法可以将儿童指导语音(CDS)作为输入,并在连续的发展水平上产生学习者的输出(LO)。我们可以将这种方法称为输入-输出(I-O)建模。在其最简单的形式中,I-O建模倾向于将语言学习视为一个紧急的,数据驱动的过程。然而,在同一个计算框架中,有足够的空间来精确地陈述非涌现主义的先天约束、参数、原则和共性的运作。此外,复杂的功能的情景背景下,社会理解和最近的对话历史,原则上,可以量化和编码的输入功能。更一般地说,只要语言习得装置(LAD; Chomsky 1965)的所有这些部分都被完全指定,
This special issue showcases recent work on the computational modeling of child language acquisition. Together, these nine papers provide testimony to the scientific value of corpus-driven computational modeling. Each of the papers presents a clear, mechanistic model that can be tested, refined or rejected on the basis of publicly available data and/or replicable experiments. However, for many readers, the multiple formalisms and technicalities involved in this type of work can serve as barriers to evaluating the nature of the contributions being made. Therefore, it is my goal in this introduction to summarize what I see as the important take-home messages delivered by each of the nine projects. Hopefully, this overview will encourage the reader to turn then to the details of each of the nine contributions. Of the nine papers included here, eight base their analysis on the corpora of spontaneous adult–child interactions made available through the Child Language Data Exchange System (CHILDES). These CHILDES corpora provide two types of information crucial to the modeling enterprise. First, they document in detail the naturalistic development of language in the child. Second, these corpora provide a good sampling of the adult speech that serves as input to the child’s language learning mechanisms. Given these two empirical bases, the job of the computational modeler is to determine a set of algorithms that can take the child-directed speech (CDS) as input and produce the learner’s output (LO) at successive developmental levels. We can refer to this approach as input–output (I–O) modeling. In its simplest form, I–O modeling tends to view language learning as an emergent, data-driven process. However, there is ample room within this same computational framework for precise statements regarding the operation of non-emergentist innate constraints, parameters, principles and universals. Furthermore, complex features of situational context, social understandings and recent dialog history can, in principle, be quantified and coded as features of the input. More generally, as long as long as all these pieces of the Language Acquisition Device (LAD; Chomsky 1965) are fully specified, all