One Billion Audio Sounds from GPU-Enabled Modular Synthesis

One Billion Audio Sounds from GPU-Enabled Modular Synthesis
复制标题

DOI:
10.23919/dafx51585.2021.9768246
复制
发表时间:
2021-04
期刊:
2021 24th International Conference on Digital Audio Effects (DAFx)
影响因子:
--
通讯作者:
Joseph P. Turian;Jordie Shier;G. Tzanetakis;K. McNally;Max Henry
Joseph P. Turian;Jordie Shier;G. Tzanetakis;K. McNally;Max Henry
中科院分区:
其他
文献类型:
--
作者:
Joseph P. Turian;Jordie Shier;G. Tzanetakis;K. McNally;Max Henry

文献摘要

被引文献

相似文献

我们发布了一个多模态音频语料库,由10亿个4秒合成声音组成,并与用于生成它们的合成参数配对。该数据集比文献中的任何音频数据集都大100倍。我们还介绍了torchsynth,这是一个开源的模块化合成器,它在单个GPU上以比实时(714MHz)快16200倍的速度动态生成synth 1B1样本。最后,我们发布了两个新的音频数据集:FM合成音色和减法合成音高。使用这些数据集,我们展示了新的基于排名的评估标准,现有的音频表示。最后,我们提出了一种新的方法来合成器超参数优化。
We release synth1B1, a multi-modal audio corpus consisting of 1 billion 4-second synthesized sounds, paired with the synthesis parameters used to generate them. The dataset is 100x larger than any audio dataset in the literature. We also introduce torchsynth, an open source modular synthesizer that generates the synth 1B1 samples on-the-fly at 16200x faster than real-time (714MHz) on a single GPU. Finally, we release two new audio datasets: FM synth timbre and subtractive synth pitch. Using these datasets, we demonstrate new rank-based evaluation criteria for existing audio representations. Finally, we propose a novel approach to synthesizer hyperparameter optimization.