Adversarial autoencoder for reducing nonlinear distortion

Adversarial autoencoder for reducing nonlinear distortion
复制标题

DOI:
10.23919/apsipa.2018.8659540
复制
发表时间:
2018-11
期刊:
2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
影响因子:
--
通讯作者:
Naohiro Tawara;Tetsunori Kobayashi;Masaru Fujieda;Kazuhiro Katagiri;T. Yazu;Tetsuji Ogawa
Naohiro Tawara;Tetsunori Kobayashi;Masaru Fujieda;Kazuhiro Katagiri;T. Yazu;Tetsuji Ogawa
中科院分区:
其他
文献类型:
--
作者:
Naohiro Tawara;Tetsunori Kobayashi;Masaru Fujieda;Kazuhiro Katagiri;T. Yazu;Tetsuji Ogawa

文献摘要

相似文献

针对时频掩蔽引起的非线性失真,提出了一种基于生成对抗网络的后置滤波方法。Tf掩蔽是用于衰减干扰声音的强大框架,但它会产生令人不快的语音失真(例如,音乐噪声)。基于GaN的自动编码器最近被证明是有效的单通道语音增强技术,然而,将该技术用于TF掩蔽的后处理并不能帮助降低非线性失真,因为在TF掩蔽后丢失了一些TF分量。此外,使用自动编码器很难嵌入丢失的信息。为了恢复这种丢失的分量,包括目标源分量的辅助参考信号被与增强信号级联,然后被用作基于GaN的自动编码器的输入。实验结果表明,本文提出的后置滤波算法在语音质量上比传统的TF掩蔽算法有明显的改善。
A novel post-filtering method using generative adversarial networks (GANs) is proposed to correct the effect of a nonlinear distortion caused by time-frequency (TF) masking. TF masking is a powerful framework for attenuating interfering sounds, but it can yield an unpleasant distortion of speech (e.g., a musical noise). A GAN-based autoencoder was recently shown to be effective for single-channel speech enhancement, however, using this technique for the post-processing of TF masking cannot help in nonlinear distortion reduction because some TF components are missing after TF-masking. Furthermore, the missing information is difficult embed using an autoencoder. In order to recover such missing components, an auxiliary reference signal that includes the target source components is concatenated with an enhanced signal, is then used as the input to the GAN-based autoencoder. Experimental comparisons show that the proposed post-filtering yields improvements in speech quality over TF-masking.