Self-Supervised Contrastive Learning for Singing Voices

Self-Supervised Contrastive Learning for Singing Voices
复制标题

歌声的自我监督对比学习

DOI:
10.1109/taslp.2022.3169627
复制
发表时间:
2022
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Masataka Goto
Masataka Goto
中科院分区:
--
文献类型:
--
作者:
Hiromu Yakura;Kento Watanabe;Masataka Goto

文献摘要

被引文献

相似文献

本研究引入自我监督的对比学习来获取歌唱声音的特征表征。为了在无监督的情况下获得稳健的表示,规则的自我监督对比学习训练神经网络使样本的特征表示接近其计算转换后的版本。同样,考虑到歌声的性质,我们采用了两种变换-音调变换和时间拉伸。然而,我们反过来使用它们:我们训练网络来推开转换后的版本的表示。然后,这些网络试图区分没有时间伸展的音调变化引起的音色变化,以及没有音高变化的音调变化引起的歌唱表达的变化。因此,习得的表征变得注重声音的音色和歌唱表达。这一点通过歌手识别任务得到了证实,在该任务中,我们训练了一个分类器来学习500名歌手的相应歌手标签与特征表示之间的关系。结果表明,与未经变换获得特征表示的情况(TOP-1准确率:53.96%)相比,所采用的变换使分类器的分类准确率提高了9.12%(TOP-1准确率:63.08%)。此外,通过改变变换的合并方式,所提出的方法可以扩展到获得关注声音音色或歌唱表达而不关注另一个的特征表示。我们特别探讨了这种以音色或歌唱表情为导向的特征表征相对于歌曲流派、歌手性别和发声技巧的特点,并证实它们成功地捕捉到了歌唱声音的不同方面。
This study introduces self-supervised contrastive learning to acquire feature representations of singing voices. To acquire robust representations in an unsupervised manner, regular self-supervised contrastive learning trains neural networks to make the feature representation of a sample close to those of its computationally transformed versions. Similarly, we employ two transformations—pitch shifting and time stretching—considering the nature of singing voices. Nevertheless, we use them reversely: we train networks to push away representations of the transformed versions. The networks then attempt to discriminate changes in vocal timbres introduced by pitch shifting without time stretching and those in singing expressions introduced by time stretching without pitch shifting. Consequently, the acquired representations become attentive to vocal timbre and singing expression. This was confirmed through a singer identification task, where we trained a classifier to learn the relationship between the feature representations to the corresponding singer labels of 500 singers. As a result, the employed transformations helped the classifier improve the classification accuracy by 9.12% (top-1 accuracy: 63.08%) compared with the case where the feature representations fed to the classifier were acquired without the transformations (top-1 accuracy: 53.96%). Furthermore, the proposed approach can be extended to acquire feature representations attentive to either vocal timbre or singing expression but not to the other by changing how the transformations are incorporated. We particularly explored the characteristics of such vocal timbre- or singing expression-oriented feature representations against song genre, singer gender, and vocal technique, and confirmed that they successfully capture different aspects of singing voices.