The Meta-Pi network: connectionist rapid adaptation for high-performance multi-speaker phoneme recognition

The Meta-Pi network: connectionist rapid adaptation for high-performance multi-speaker phoneme recognition
复制标题

Meta-Pi 网络:连接主义快速适应高性能多说话人音素识别

DOI:
10.1109/icassp.1990.115564
复制
发表时间:
1990
期刊:
International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
A. Waibel
A. Waibel
中科院分区:
--
文献类型:
--
作者:
J. Hampshire;A. Waibel

文献摘要

被引文献

相似文献

提出了一种基于多网络时延神经网络(TDNN)的连接主义结构,该结构允许在98.4%的说话人相关识别率下进行多说话人音素辨别(/B,d,g/)。整个网络门控在个体说话者上训练的模块的音素决定,以形成其总体分类决定。通过动态地适应输入语音并专注于特定于说话者的模块的组合,该网络的性能优于在所有六个说话者的语音上训练的单个TDNN(95.9%)。为了训练这个网络,开发了一种称为Meta-Pi连接的乘法连接形式。它说明了如何Mega-Pi范例实现一个动态自适应贝叶斯MAP分类器。它学习-没有监督-认识到一个特定的发言者(99.8%)使用其他发言者的内部模型的动态组合。Meta-Pi模型是连接主义语音识别系统的可行基础,可以快速适应新的说话者和不同的说话者方言。&lt;<ETX>&gt;
A multinetwork time-delay-neural-network (TDNN)-based connectionist architecture that allows multispeaker phoneme discrimination (/b,d,g/) to be performed at the speaker-dependent recognition rate of 98.4% is presented. The overall network gates the phonemic decisions of modules trained on individual speakers to form its overall classification decision. By dynamically adapting to the input speech and focusing on a combination of speaker-specific modules, the network outperforms a single TDNN trained on the speech of all six speakers (95.9%). To train this network a form of multiplicative connection called the Meta-Pi connection is developed. It is illustrated how the Mega-Pi paradigm implements a dynamically adaptive Bayesian MAP classifier. It learns-without supervision-to recognize the speech of one particular speaker (99.8%) using a dynamic combination of internal models of other speakers exclusively. The Meta-Pi model is a viable basis for a connectionist speech recognition system that can rapidly adapt to new speakers and varying speaker dialects.<<ETX>>