A mapping model of spectral tilt in normal-to-Lombard speech conversion for intelligibility enhancement

A mapping model of spectral tilt in normal-to-Lombard speech conversion for intelligibility enhancement
复制标题

用于增强清晰度的普通到伦巴第语音转换中的频谱倾斜映射模型

DOI:
10.1007/s11042-020-08838-1
复制
发表时间:
2020-03
影响因子:
3.6
通讯作者:
Wang Xiaochen
Wang Xiaochen
中科院分区:
计算机科学4区
文献类型:
--
作者:
Li Gang;Hu Ruimin;Zhang Rui;Wang Xiaochen

文献摘要

参考文献

相似文献

环境噪声降低了听电话时的语音清晰度。虽然手机有干净的信号源,但听者还是很难获得信息。可懂度增强(IENH)是一种用于在噪声环境中呈现清晰语音的感知增强技术。本研究受Lombard反射的启发,通过正常到Lombard语音转换来实现IENH。在这个转换过程中,关键是将正常语音(正常风格)的频谱倾斜映射到伦巴第语音(伦巴第风格)。对于映射的光谱倾斜,我们提出了一个映射模型相结合的线性预测为基础的映射网络和倾斜修改。与以前的研究相比,我们使用深度神经网络(DNN)而不是基于高斯的模型进行高维映射,并创造性地添加了倾斜修改模块,以进一步减少共振峰幅度的映射误差。在本文中,我们使用AVS-M编解码器和两个数据集作为基准平台。评价结果表明,我们的方法得到了更好的结果比参考方法在客观和主观的实验。
Environmental noise degrades the speech intelligibility when listening to the phone. Although the phone has a clean signal source, it is still difficult for the listener to get information. Intelligibility enhancement (IENH) is a type of perceptual enhancement technique for clean speech rendered in noisy environments. This study focuses on IENH by normal-to-Lombard speech conversion, which is inspired by Lombard reflex. In this conversion process, the key point is to map the spectral tilt from the normal speech (normal style) to the Lombard speech (Lombard style). For mapping the spectral tilt, we propose a mapping model combining linear-prediction-based mapping networks and tilt modification. Compared with previous studies, we use deep neural networks (DNNs) instead of Gaussian-based models for higher dimensional mapping, and inventively add a tilt modification module to reduce the mapping errors of formant magnitudes further. In this paper, we use AVS-M codec and two datasets as the benchmark platform. The valuation shows that our method gets better results than reference methods in both objective and subjective experiments.
DOI: 10.1109/icassp.1991.150351
发表时间: 1991-04
期刊: [Proceedings] ICASSP 91: 1991 International Conference on Acoustics, Speech, and Signal Processing
影响因子: --
作者:
J. Junqua
通讯作者: J. Junqua
DOI: --
发表时间: 2010-03
期刊: --
影响因子: --
作者:
L. Rabiner;R. Schafer
通讯作者: L. Rabiner;R. Schafer
DOI: 10.1109/icassp.2014.6854483
发表时间: 2014-05
期刊: 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
M. Koutsogiannaki;Y. Stylianou
通讯作者: M. Koutsogiannaki;Y. Stylianou
DOI: 10.1145/2499788.2499839
发表时间: 2013-08
期刊: --
影响因子: --
作者:
Xiaochen Wang;Y. Wang;B. Hang
通讯作者: Xiaochen Wang;Y. Wang;B. Hang
DOI: 10.21437/interspeech.2013-769
发表时间: 2013
期刊: --
影响因子: --
作者:
H. Schepker;J. Rennies;S. Doclo
通讯作者: H. Schepker;J. Rennies;S. Doclo