Simple and artefact-free spectral modifications for enhancing the intelligibility of casual speech

Simple and artefact-free spectral modifications for enhancing the intelligibility of casual speech
复制标题

DOI:
10.1109/icassp.2014.6854483
复制
发表时间:
2014-05
期刊:
2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
M. Koutsogiannaki;Y. Stylianou
M. Koutsogiannaki;Y. Stylianou
中科院分区:
其他
文献类型:
--
作者:
M. Koutsogiannaki;Y. Stylianou

文献摘要

被引文献

相似文献

在本文中,解决了修改随意语音以达到清晰语音的可懂度水平的问题。与其他研究不同,在这项工作中,对随意语音的修改同时考虑了清晰度和语音质量。为了实现这一目标,作者将重点放在受清晰语音启发的类人修改上。对清晰和随意的语音进行的声学分析揭示了两种说话风格之间特定频段的能量差异。然后,使用一种简单的方法来增强随意语音的这些频率区域。所提出的方法称为混合滤波,使用多频带滤波方案来隔离这些频带的信息,然后将该信息添加到原始信号中。我们的方法在清晰度和质量方面与未经修改的随意语音以及高度可理解的频谱修改技术(即频谱整形和动态范围压缩(SSDRC))进行了比较。与主观清晰度分数高度相关的两种不同的客观测量用于估计清晰度,而为了评估质量,进行偏好听力测试。结果表明,混合过滤技术提高了随意语音的清晰度,同时保持了其质量。另一方面,虽然 SSDRC 在清晰度方面表现出色,但它显着降低了随意语音的质量。
In this paper, the problem of modifying casual speech to reach the intelligibility level of clear speech is addressed. Unlike other studies, in this work modifications on casual speech both consider intelligibility and speech quality. To achieve this, the authors focus on human-like modifications inspired by clear speech. An acoustic analysis performed on clear and casual speech reveals energy differences on specific frequency bands between the two speaking styles. Then, a simple method is used to boost these frequency regions on casual speech. The proposed method, called mix-filtering, uses a multi-band filtering scheme to isolate the information of these frequency bands and then, add this information to the original signal. Our method is compared in terms of intelligibility and quality with unmodified casual speech and with a highly intelligible spectral modification technique, namely the Spectral Shaping and Dynamic Range Compression (SSDRC). Two different objective measures that are highly correlated with subjective intelligibility scores are used for estimating the intelligibility, whereas for evaluating the quality, preference listening tests are performed. Results show that the mix-filtering technique increases the intelligibility of casual speech while maintains its quality. On the other hand, while SSDRC outperforms on intelligibility, it degrades significantly the quality of casual speech.