Reconstruction of Phonated Speech from Whispers Using Formant-Derived Plausible Pitch Modulation

Reconstruction of Phonated Speech from Whispers Using Formant-Derived Plausible Pitch Modulation
复制标题

使用共振峰衍生的合理音高调制从耳语中重建发声语音

DOI:
10.1145/2737724
复制
发表时间:
2015
期刊:
ACM Transactions on Accessible Computing (TACCESS)
影响因子:
--
通讯作者:
Yan Song
Yan Song
中科院分区:
--
文献类型:
--
作者:
I. Mcloughlin;H. Sharifzadeh;Su;Jingjie Li;Yan Song

文献摘要

参考文献

被引文献

相似文献

对大多数人来说,耳语是一种自然的、不发音的、次要的语言交流方式。然而,对于一些声音产生机制受损的说话者来说,这是主要的交流机制,比如部分喉部切除术,以及那些通常在手术或喉部损伤后规定的声音休息。不像大多数人,他们选择什么时候轻声说话,什么时候不轻声说话,这些人可能别无选择,只能依靠轻声说话来进行日常的声音交流。尽管大多数说话者有时会低声说话,有些说话者只能低声说话,但今天的大多数计算语音技术系统都假设或需要语音发音。这篇文章考虑将耳语转化为自然发音的语音,作为一种非侵入性的假肢帮助那些只能耳语的声音障碍人士。作为副产品,这项技术对选择耳语的正常人也很有用。语音重建系统可以分为需要训练的系统和不需要训练的系统。在后者中,最近的参数重建框架进行了探索,然后通过加权形成峰差异的合理音高的改进估计。改进的重建框架,提出的共振峰衍生的人工音调调制,通过主观和客观的比较测试和最先进的替代方案进行了验证。
Whispering is a natural, unphonated, secondary aspect of speech communications for most people. However, it is the primary mechanism of communications for some speakers who have impaired voice production mechanisms, such as partial laryngectomees, as well as for those prescribed voice rest, which often follows surgery or damage to the larynx. Unlike most people, who choose when to whisper and when not to, these speakers may have little choice but to rely on whispers for much of their daily vocal interaction. Even though most speakers will whisper at times, and some speakers can only whisper, the majority of today’s computational speech technology systems assume or require phonated speech. This article considers conversion of whispers into natural-sounding phonated speech as a noninvasive prosthetic aid for people with voice impairments who can only whisper. As a by-product, the technique is also useful for unimpaired speakers who choose to whisper. Speech reconstruction systems can be classified into those requiring training and those that do not. Among the latter, a recent parametric reconstruction framework is explored and then enhanced through a refined estimation of plausible pitch from weighted formant differences. The improved reconstruction framework, with proposed formant-derived artificial pitch modulation, is validated through subjective and objective comparison tests alongside state-of-the-art alternatives.
使用高斯混合模型进行 NAM 到语音的转换
DOI: --
发表时间: 2005
期刊: Proceedings of Interspeech 2005
影响因子: --
作者:
Tomoki Toda;Kiyohiro Shikano
通讯作者: Kiyohiro Shikano
通过统计语音转换改善身体传输的清音语音
DOI: --
发表时间: 2006
期刊:
影响因子: --
作者:
Mikihiro Nakagiri;Tomoki Toda;Hideki Kashioka;Kiyohiro Shikano
通讯作者: Kiyohiro Shikano