Flow Synthesizer: Universal Audio Synthesizer Control with Normalizing Flows

Flow Synthesizer: Universal Audio Synthesizer Control with Normalizing Flows
复制标题

流合成器:具有标准化流的通用音频合成器控制

DOI:
10.3390/app10010302
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Axel Chemla
Axel Chemla
中科院分区:
--
文献类型:
--
作者:
P. Esling;Naotake Masuda;Adrien Bardet;R. Despres;Axel Chemla

文献摘要

被引文献

相似文献

无处不在的声音合成器已经重塑了现代音乐制作,新的音乐流派现在甚至完全由它们的使用来定义。然而,不断增加的复杂性和参数在现代合成器的数量使他们极其难以掌握。因此,方法的发展允许轻松地创建和探索与合成器是一个至关重要的需要。最近,我们介绍了一种新的音频合成器控制公式,该公式基于学习合成器能力的有组织潜在音频空间,同时构造到其参数空间的可逆映射。我们表明,该公式允许在单个模型中同时处理自动参数推理、宏观控制学习和基于音频的预设探索。我们证明了这个公式可以通过依赖于变分自编码器(VAE)和规范化流(NF)有效地解决。在本文中,我们通过在更大的参数集上评估我们的建议来扩展我们的结果,并展示了它在参数推理和针对各种基线模型的音频重建方面的优势。此外,我们引入了解纠缠流,它允许学习两个独立潜在空间之间的可逆映射,同时通过将目标分割为部分密度评估来指导某些潜在维度的组织以匹配目标变化因子。我们表明,该模型将音频变化的主要因素分解为潜在维度,这些潜在维度可以直接用作宏观参数。我们还表明,我们的模型能够学习合成器的语义控制,同时平滑地映射到其参数。最后,我们在实时Max4Live设备中介绍了我们模型的开源实现,该设备可随时用于评估我们提案的创造性应用。
The ubiquity of sound synthesizers has reshaped modern music production, and novel music genres are now sometimes even entirely defined by their use. However, the increasing complexity and number of parameters in modern synthesizers make them extremely hard to master. Hence, the development of methods allowing to easily create and explore with synthesizers is a crucial need. Recently, we introduced a novel formulation of audio synthesizer control based on learning an organized latent audio space of the synthesizer’s capabilities, while constructing an invertible mapping to the space of its parameters. We showed that this formulation allows to simultaneously address automatic parameters inference, macro-control learning, and audio-based preset exploration within a single model. We showed that this formulation can be efficiently addressed by relying on Variational Auto-Encoders (VAE) and Normalizing Flows (NF). In this paper, we extend our results by evaluating our proposal on larger sets of parameters and show its superiority in both parameter inference and audio reconstruction against various baseline models. Furthermore, we introduce disentangling flows, which allow to learn the invertible mapping between two separate latent spaces, while steering the organization of some latent dimensions to match target variation factors by splitting the objective as partial density evaluation. We show that the model disentangles the major factors of audio variations as latent dimensions, which can be directly used as macro-parameters. We also show that our model is able to learn semantic controls of a synthesizer, while smoothly mapping to its parameters. Finally, we introduce an open-source implementation of our models inside a real-time Max4Live device that is readily available to evaluate creative applications of our proposal.