GACELA: A Generative Adversarial Context Encoder for Long Audio Inpainting of Music

GACELA: A Generative Adversarial Context Encoder for Long Audio Inpainting of Music
复制标题

GACELA:用于音乐长音频修复的生成对抗性上下文编码器

DOI:
--
复制
发表时间:
2020
期刊:
IEEE Journal on Selected Topics in Signal Processing
影响因子:
--
通讯作者:
Nathanael Perraudin
Nathanael Perraudin
中科院分区:
--
文献类型:
--
作者:
Andrés Marafioti;P. Majdak;N. Holighaus;Nathanael Perraudin

文献摘要

参考文献

被引文献

相似文献

在本文中,我们介绍了GACELA,一种条件生成对抗网络(cGAN),旨在恢复持续时间在数百毫秒到几秒之间的丢失音频数据,即,来执行长间隙音频修复。虽然以前的工作要么解决了较短的间隙,要么依赖于样本,从其他信号部分复制可用的信息,GACELA解决了两个方面的长间隙的修复。首先,它认为不同的时间尺度的音频信息依赖于五个平行的鉴别器,增加分辨率的感受野。第二,它不仅取决于围绕差距的可用信息,即,上下文,但也对cGAN的潜在变量。这解决了这种长间隙的音频修复的固有多模态,同时为用户提供不同的修复选项。GACELA在不同的复杂度和不同的间隙持续时间从375到1500毫秒的音乐信号的听力测试中进行了评估。在实验室条件下,我们的受试者通常能够检测到修复。然而,修复后伪影的严重度被评定为不干扰和轻度干扰之间。GACELA代表了一个框架,能够集成未来的改进,如处理更多的音乐相关的功能或明确的音乐功能。我们的软件和经过训练的模型,辅以指导性的示例,可在线获得。
In this article, we introduce GACELA, a conditional generative adversarial network (cGAN) designed to restore missing audio data with durations ranging between hundreds of milliseconds and a few seconds, i.e., to perform long-gap audio inpainting. While previous work either addressed shorter gaps or relied on exemplars by copying available information from other signal parts, GACELA addresses the inpainting of long gaps in two aspects. First, it considers various time scales of audio information by relying on five parallel discriminators with increasing resolution of receptive fields. Second, it is conditioned not only on the available information surrounding the gap, i.e., the context, but also on the latent variable of the cGAN. This addresses the inherent multi-modality of audio inpainting for such long gaps while providing the user with different inpainting options. GACELA was evaluated in listening tests on music signals of varying complexity and varying gap durations from 375 to 1500 ms. Under laboratory conditions, our subjects were often able to detect the inpainting. However, the severity of the inpainted artifacts was rated between not disturbing and mildly disturbing. GACELA represents a framework capable of integrating future improvements such as processing of more auditory-related features or explicit musical features. Our software and trained models, complemented by instructive examples, are available online.
DOI: 10.1109/icassp.2011.5946407
发表时间: 2011-05
期刊: 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
A. Adler;Valentin Emiya;M. Jafari;Michael Elad;R. Gribonval;Mark D. Plumbley
通讯作者: A. Adler;Valentin Emiya;M. Jafari;Michael Elad;R. Gribonval;Mark D. Plumbley