Lyrics-to-Audio Alignment by Unsupervised Discovery of Repetitive Patterns in Vowel Acoustics

Lyrics-to-Audio Alignment by Unsupervised Discovery of Repetitive Patterns in Vowel Acoustics
复制标题

通过无监督地发现元音声学中的重复模式来实现歌词到音频的对齐

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
3.9
通讯作者:
Kyogu Lee
Kyogu Lee
中科院分区:
计算机科学3区
文献类型:
--
作者:
Sungkyun Chang;Kyogu Lee

文献摘要

被引文献

相似文献

之前大多数歌词到音频对齐的方法都使用预先开发的自动语音识别 (ASR) 系统,该系统在使语音模型适应个别歌手方面存在一些困难。以前的作品中缺少的一个重要方面是歌声中重复元音模式的自学习性,其中使用的元音部分比辅音部分更加一致。基于此,我们的系统首先通过将标准声学特征的自相似性作为输入,基于加权对称非负矩阵分解来学习元音序列的判别子空间。然后,我们利用源自最新计算机视觉技术的规范时间扭曲来找到文本和声音序列之间的最佳时空变换。对韩语和英语数据集的实验表明,在预先开发的、无监督的歌源分离之后部署此方法比其他最先进的无监督方法和现有的基于 ASR 的系统取得了更有希望的结果。
Most of the previous approaches to lyrics-to-audio alignment used a pre-developed automatic speech recognition (ASR) system that innately suffered from several difficulties to adapt the speech model to individual singers. A significant aspect missing in previous works is the self-learnability of repetitive vowel patterns in the singing voice, where the vowel part used is more consistent than the consonant part. Based on this, our system first learns a discriminative subspace of vowel sequences, based on weighted symmetric non-negative matrix factorization, by taking the self-similarity of a standard acoustic feature as an input. Then, we make use of canonical time warping, derived from a recent computer vision technique, to find an optimal spatiotemporal transformation between the text and the acoustic sequences. Experiments with Korean and English data sets showed that deploying this method after a pre-developed, unsupervised, singing source separation achieved more promising results than the other state-of-the-art unsupervised approaches and an existing ASR-based system.