A neural network model of the articulatory-acoustic forward mapping trained on recordings of articulatory parameters.

A neural network model of the articulatory-acoustic forward mapping trained on recordings of articulatory parameters.
复制标题

DOI:
10.1121/1.1715112
复制
发表时间:
2004-10
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
通讯作者:
Christopher T. Kello;D. Plaut
Christopher T. Kello;D. Plaut
中科院分区:
其他
文献类型:
--
作者:
Christopher T. Kello;D. Plaut

文献摘要

被引文献

相似文献

三个神经网络模型进行了训练的前向映射从发音位置的声学输出为一个单一的发言人的爱丁堡多通道发音语音数据库。模型参数(即,连接权重)是经由误差信号的反向传播来学习的,该误差信号由模型的声学输出与它们的声学目标之间的差生成。通过对模型的声学输出进行语音清晰度测试来评估训练模型的有效性。这些测试的结果表明,足够的语音信息被捕获的模型,以支持高达84%的单词识别率,接近92%的实际目标刺激的识别率。这些前向模型可以作为数据驱动的发音合成器的一个组成部分。这些模型还为建立一个基于真实的语音训练的口语词汇习得和语音发展模型迈出了第一步。
Three neural network models were trained on the forward mapping from articulatory positions to acoustic outputs for a single speaker of the Edinburgh multi-channel articulatory speech database. The model parameters (i.e., connection weights) were learned via the backpropagation of error signals generated by the difference between acoustic outputs of the models, and their acoustic targets. Efficacy of the trained models was assessed by subjecting the models' acoustic outputs to speech intelligibility tests. The results of these tests showed that enough phonetic information was captured by the models to support rates of word identification as high as 84%, approaching an identification rate of 92% for the actual target stimuli. These forward models could serve as one component of a data-driven articulatory synthesizer. The models also provide the first step toward building a model of spoken word acquisition and phonological development trained on real speech.