Vocal Tract Length Normalization for Speaker Independent Acoustic-to-Articulatory Speech Inversion

Vocal Tract Length Normalization for Speaker Independent Acoustic-to-Articulatory Speech Inversion
复制标题

用于独立于说话人的声学到发音语音反转的声带长度标准化

DOI:
--
复制
发表时间:
2016
期刊:
Interspeech
影响因子:
--
通讯作者:
C. Espy
C. Espy
中科院分区:
--
文献类型:
--
作者:
G. Sivaraman;V. Mitra;Hosung Nam;M. Tiede;C. Espy

文献摘要

参考文献

被引文献

相似文献

语音反转是一个众所周知的不适定问题,加上说话人的差异通常会使它变得更加困难。本文研究了一种声道长度归一化(VTLN)技术,将不同说话人的声学空间转换到目标说话人空间,使说话人的具体细节最小化。说话人归一化的功能,然后用来训练基于前馈神经网络的声学发音语音反演系统。声学特征被参数化为时间上下文化的梅尔频率倒谱系数,发音特征由六个声道变量(TV)轨迹表示。实验是用10个来自美国的说话人进行的。Wisc. X射线微束数据库。以每个说话人为基线训练说话人相关语音反转系统,以比较说话人无关方法的性能。对于每个目标扬声器,从其余九个扬声器的数据进行变换,使用所提出的方法和变换后的功能被用来训练语音反转系统。使用目标说话人测试集上的估计电视和实际电视之间的相关性来比较各个系统的性能。结果表明,所提出的说话人归一化方法提供了一个7%的绝对改善相关性相比,系统中没有进行说话人归一化。
Speech inversion is a well-known ill-posed problem and addition of speaker differences typically makes it even harder. This paper investigates a vocal tract length normalization (VTLN) technique to transform the acoustic space of different speakers to a target speaker space such that speaker specific details are minimized. The speaker normalized features are then used to train a feed-forward neural network based acoustic-toarticulatory speech inversion system. The acoustic features are parameterized as time-contextualized mel-frequency cepstral coefficients and the articulatory features are represented by six tract-variable (TV) trajectories. Experiments are performed with ten speakers from the U. Wisc. X-ray microbeam database. Speaker dependent speech inversion systems are trained for each speaker as baselines to compare the performance of the speaker independent approach. For each target speaker, data from the remaining nine speakers are transformed using the proposed approach and the transformed features are used to train a speech inversion system. The performances of the individual systems are compared using the correlation between the estimated and the actual TVs on the target speaker’s test set. Results show that the proposed speaker normalization approach provides a 7% absolute improvement in correlation as compared to the system where speaker normalization was not performed.
根据语音声学估计手势分数的过程。
DOI: 10.1121/1.4763545
发表时间: 2012
期刊: The Journal of the Acoustical Society of America
影响因子: --
作者:
Nam,Hosung;Mitra,Vikramjit;Tiede,Mark;Hasegawa-Johnson,Mark;Espy-Wilson,Carol;Saltzman,Elliot;Goldstein,Louis
通讯作者: Goldstein,Louis