Glove-talk II - a neural-network interface which maps gestures to parallel formant speech synthesizer controls

Glove-talk II - a neural-network interface which maps gestures to parallel formant speech synthesizer controls
复制标题

Glove-talk II - 一种神经网络接口,可将手势映射到并行共振峰语音合成器控件

DOI:
10.1109/72.623199
复制
发表时间:
1997
影响因子:
--
通讯作者:
Geoffrey E. Hinton
Geoffrey E. Hinton
中科院分区:
--
文献类型:
--
作者:
S. Fels;Geoffrey E. Hinton

文献摘要

被引文献

相似文献

Glove-Talk II是一个通过自适应界面将手势转换为语音的系统。手势连续映射到并行共振峰语音合成器的10个控制参数。这种映射允许手充当人造声道,在真实的时间内产生语音。除了直接控制基频和音量外,这还提供了无限的词汇量。目前,Glove-Talk II的最佳版本使用几个输入设备,一个并行共振峰语音合成器和三个神经网络。通过使用门控网络对元音和辅音神经网络的输出进行加权,将手势到语音任务分为元音和辅音产生。门控网络和辅音网络是用来自用户的示例来训练的。元音网络实现了手的位置和元音声音之间的固定用户定义的关系,不需要来自用户的任何训练示例。音量、基频和塞音都是通过输入设备的固定映射产生的。使用手套对话II,受试者可以慢慢地说话,但比文本到语音合成器更自然的音调变化。
Glove-Talk II is a system which translates hand gestures to speech through an adaptive interface. Hand gestures are mapped continuously to ten control parameters of a parallel formant speech synthesizer. The mapping allows the hand to act as an artificial vocal tract that produces speech in real time. This gives an unlimited vocabulary in addition to direct control of fundamental frequency and volume. Currently, the best version of Glove-Talk II uses several input devices, a parallel formant speech synthesizer, and three neural networks. The gesture-to-speech task is divided into vowel and consonant production by using a gating network to weight the outputs of a vowel and a consonant neural network. The gating network and the consonant network are trained with examples from the user. The vowel network implements a fixed user-defined relationship between hand position and vowel sound and does not require any training examples from the user. Volume, fundamental frequency, and stop consonants are produced with a fixed mapping from the input devices. With Glove-Talk II, the subject can speak slowly but with far more natural sounding pitch variations than a text-to-speech synthesizer.