Direct Speech Reconstruction From Articulatory Sensor Data by Machine Learning

Direct Speech Reconstruction From Articulatory Sensor Data by Machine Learning
复制标题

DOI:
10.1109/taslp.2017.2757263
复制
发表时间:
2017-12-01
影响因子:
5.4
通讯作者:
Holdsworth, Ed
Holdsworth, Ed
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gonzalez, Jose A.;Cheah, Lam A.;Holdsworth, Ed

文献摘要

被引文献

相似文献

本文描述了一种由发音器运动产生语音声学的技术。我们的动机是帮助那些喉切除术后不能说话的人,喉切除术在西方世界每年进行数万次。我们的方法来感知发音器的运动,永久磁性发音仪,依赖于小的,不显眼的磁铁附着在嘴唇和舌头。由磁铁运动引起的磁场变化被感知,并形成一个过程的输入,该过程被训练来估计语音声学。在这里报告的实验中,这种“直接合成”技术是为普通扬声器开发的,使用胶合磁铁,使我们能够使用平行传感器和声学数据进行训练。我们描述了基于高斯混合模型、深度神经网络和循环神经网络(rnn)的三种机器学习技术。我们通过客观声学失真测量和主观听力测试来评估我们的技术,这些测试来自小说(CMU北极语料库)的口语句子。我们的研究结果表明,表现最好的技术是双向RNN (BiRNN),它利用过去和未来的背景来从传感器数据中预测声学。birnn不适合实时合成,但固定滞后rnn给出了类似的结果,并且因为它们只展望未来的一小段路,克服了这个问题。听力测试表明,通过这种方法产生的语音具有自然的质量,保留了说话者的身份。此外,我们在具有挑战性的CMU北极材料上获得了高达92%的清晰度。据我们所知,这些是无声语音系统在没有限制词汇和不引人注目的设备的情况下获得的最佳结果,该设备可以提供接近实时的音频。这项工作有望带来一种技术,真正让喉部被切除的人恢复声音。
This paper describes a technique that generates speech acoustics from articulator movements. Our motivation is to help people who can no longer speak following laryngectomy, a procedure that is carried out tens of thousands of times per year in the Western world. Our method for sensing articulator movement, permanent magnetic articulography, relies on small, unobtrusive magnets attached to the lips and tongue. Changes in magnetic field caused by magnet movements are sensed and form the input to a process that is trained to estimate speech acoustics. In the experiments reported here this "Direct Synthesis" technique is developed for normal speakers, with glued-on magnets, allowing us to train with parallel sensor and acoustic data. We describe three machine learning techniques for this task, based on Gaussian mixture models, deep neural networks, and recurrent neural networks (RNNs). We evaluate our techniques with objective acoustic distortion measures and subjective listening tests over spoken sentences read from novels (the CMU Arctic corpus). Our results show that the best performing technique is a bidirectional RNN (BiRNN), which employs both past and future contexts to predict the acoustics from the sensor data. BiRNNs are not suitable for synthesis in real time but fixed-lag RNNs give similar results and, because they only look a little way into the future, overcome this problem. Listening tests show that the speech produced by this method has a natural quality that preserves the identity of the speaker. Furthermore, we obtain up to 92% intelligibility on the challenging CMU Arctic material. To our knowledge, these are the best results obtained for a silent-speech system without a restricted vocabulary and with an unobtrusive device that delivers audio in close to real time. This work promises to lead to a technology that truly will give people whose larynx has been removed their voices back.