Vocal Tract Length Normalization for Speaker Independent Acoustic-to-Articulatory Speech Inversion
Vocal Tract Length Normalization for Speaker Independent Acoustic-to-Articulatory Speech Inversion
复制标题
用于独立于说话人的声学到发音语音反转的声带长度标准化
DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
C. Espy
中科院分区:
文献类型:
--
作者:
G. Sivaraman;V. Mitra;Hosung Nam;M. Tiede;C. Espy
Speech inversion is a well-known ill-posed problem and addition of speaker differences typically makes it even harder. This paper investigates a vocal tract length normalization (VTLN) technique to transform the acoustic space of different speakers to a target speaker space such that speaker specific details are minimized. The speaker normalized features are then used to train a feed-forward neural network based acoustic-toarticulatory speech inversion system. The acoustic features are parameterized as time-contextualized mel-frequency cepstral coefficients and the articulatory features are represented by six tract-variable (TV) trajectories. Experiments are performed with ten speakers from the U. Wisc. X-ray microbeam database. Speaker dependent speech inversion systems are trained for each speaker as baselines to compare the performance of the speaker independent approach. For each target speaker, data from the remaining nine speakers are transformed using the proposed approach and the transformed features are used to train a speech inversion system. The performances of the individual systems are compared using the correlation between the estimated and the actual TVs on the target speaker’s test set. Results show that the proposed speaker normalization approach provides a 7% absolute improvement in correlation as compared to the system where speaker normalization was not performed.
DOI:
10.1121/1.4763545
发表时间:
2012
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
作者:
Nam,Hosung;Mitra,Vikramjit;Tiede,Mark;Hasegawa-Johnson,Mark;Espy-Wilson,Carol;Saltzman,Elliot;Goldstein,Louis
通讯作者:
Goldstein,Louis