Pareto-Optimized Non-Negative Matrix Factorization Approach to the Cleaning of Alaryngeal Speech Signals.

Pareto-Optimized Non-Negative Matrix Factorization Approach to the Cleaning of Alaryngeal Speech Signals.
复制标题

DOI:
10.3390/cancers15143644
复制
发表时间:
2023-07-16
期刊:
影响因子:
5.2
通讯作者:
Uloza, Virgilijus
Uloza, Virgilijus
中科院分区:
医学2区
文献类型:
--
作者:
Maskeliunas, Rytis;Damasevicius, Robertas;Kulikajevas, Audrius;Pribuisis, Kipras;Ulozaite-Staniene, Nora;Uloza, Virgilijus

文献摘要

参考文献

相似文献

本文介绍了一种通过将Pareto优化深度学习与非负矩阵分解(NMF)相结合来清理受损语音的新方法。该方法有效地降低了受损语音中的噪声,同时保持了所需的语音质量。该方法包括计算有噪声的语音片段的频谱图,确定噪声阈值,计算噪声到信号掩码,并对其进行平滑以避免突然的过渡。使用Pareto优化的NMF,修改后的频谱图被分解成基函数和权重,允许重建干净的语音频谱图。最终结果是通过反转干净的语音频谱图实现的噪声降低的波形。实验结果验证了该方法在去除无喉语音信号中的有效性,表明了其潜在的实际应用价值。受损语音的清除问题对于语音识别、电信和辅助技术等各种应用至关重要。在本文中,我们提出了一种新的方法,将Pareto优化的深度学习与非负矩阵分解(NMF)相结合,以有效地降低受损语音信号中的噪声,同时保持所需语音的质量。我们的方法首先计算一个嘈杂的语音片段的频谱图和提取频率统计。然后基于期望的噪声灵敏度确定阈值,并且计算噪声-信号掩模。该掩模被平滑以避免噪声水平的突然转变,并且通过将平滑掩模应用于信号频谱图来获得修改的频谱图。然后,我们使用帕累托优化的NMF将修改后的谱图分解为基函数和相应的权重,用于重建干净的语音谱图。通过对干净的语音频谱图进行反转来获得最终的噪声降低的波形。我们提出的方法通过在深度学习模型中利用帕累托优化来实现各种目标之间的平衡,例如噪声抑制,语音质量保持和计算效率。实验结果表明,我们的方法在清理喉音语音信号的有效性,使其成为一个有前途的解决方案,为各种现实世界的应用。
This paper introduces a new method for cleaning impaired speech by combining Pareto-optimized deep learning with Non-negative Matrix Factorization (NMF). The approach effectively reduces noise in impaired speech while preserving the desired speech quality. The method involves calculating the spectrogram of a noisy voice clip, determining a noise threshold, computing a noise-to-signal mask, and smoothing it to avoid abrupt transitions. Using a Pareto-optimized NMF, the modified spectrogram is decomposed into basis functions and weights, allowing for reconstruction of the clean speech spectrogram. The final result is a noise-reduced waveform achieved by inverting the clean speech spectrogram. Experimental results validate the method’s effectiveness in cleaning alaryngeal speech signals, indicating its potential for real-world applications. The problem of cleaning impaired speech is crucial for various applications such as speech recognition, telecommunication, and assistive technologies. In this paper, we propose a novel approach that combines Pareto-optimized deep learning with non-negative matrix factorization (NMF) to effectively reduce noise in impaired speech signals while preserving the quality of the desired speech. Our method begins by calculating the spectrogram of a noisy voice clip and extracting frequency statistics. A threshold is then determined based on the desired noise sensitivity, and a noise-to-signal mask is computed. This mask is smoothed to avoid abrupt transitions in noise levels, and the modified spectrogram is obtained by applying the smoothed mask to the signal spectrogram. We then employ a Pareto-optimized NMF to decompose the modified spectrogram into basis functions and corresponding weights, which are used to reconstruct the clean speech spectrogram. The final noise-reduced waveform is obtained by inverting the clean speech spectrogram. Our proposed method achieves a balance between various objectives, such as noise suppression, speech quality preservation, and computational efficiency, by leveraging Pareto optimization in the deep learning model. The experimental results demonstrate the effectiveness of our approach in cleaning alaryngeal speech signals, making it a promising solution for various real-world applications.
DOI: 10.1002/hed.24918
发表时间: 2017-12
期刊: Head & neck
影响因子: --
作者:
Birkeland AC;Beesley L;Bellile E;Rosko AJ;Hoesli R;Chinn SB;Shuman AG;Prince ME;Wolf GT;Bradford CR;Brenner JC;Spector ME
通讯作者: Spector ME
优化依赖说话者的特征提取参数,以改善构音障碍患者的自动语音识别性能。
DOI: 10.3390/s21196460
发表时间: 2021-09-27
期刊: Sensors (Basel, Switzerland)
影响因子: --
作者:
Marini M;Vanello N;Fanucci L
通讯作者: Fanucci L
DOI: 10.1186/s12938-021-00915-2
发表时间: 2021-08-03
影响因子: 3.9
作者:
Fu J;Yang S;He F;He L;Li Y;Zhang J;Xiong X
通讯作者: Xiong X
DOI: 10.6004/jnccn.2022.0016
发表时间: 2022-03-01
影响因子: 13.4
作者:
Caudell, Jimmy J.;Gillison, Maura L.;Darlow, Susan D.
通讯作者: Darlow, Susan D.