MusicYOLO: A Vision-Based Framework for Automatic Singing Transcription

MusicYOLO: A Vision-Based Framework for Automatic Singing Transcription
复制标题

MusicYOLO:基于视觉的自动歌唱转录框架

DOI:
10.1109/taslp.2022.3221005
复制
发表时间:
2023
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Wenqing Cheng
Wenqing Cheng
中科院分区:
其他
文献类型:
--
作者:
Xianke Wang;Bowen Tian;Weiming Yang;Wei Xu;Wenqing Cheng

文献摘要

参考文献

相似文献

自动歌唱转录(AST)是指从歌唱音频中推断起点、终点和音高的过程,在音乐信息检索中具有重要意义。大多数AST模型使用卷积神经网络来提取光谱特征并分别预测起始时刻和偏移时刻。该方法首先推导出帧级概率,然后通过后处理得到音符级的转录结果。本文提出了一个新的AST框架MusicYOLO,它直接获得音符级的转录结果。起点/终点检测基于目标检测模型YOLOX,基音标记通过谱峰搜索完成。与以前的方法相比,MusicYOLO检测音符对象,而不是孤立的开始/结束时刻,从而大大提高了转录性能。在本文建立的视唱声数据集(SSVD)上,MusicYOLO获得了84.60%的转录F1分数,这是最先进的方法。
Automatic singing transcription (AST), which refers to the process of inferring the onset, offset, and pitch from the singing audio, is of great significance in music information retrieval. Most AST models use the convolutional neural network to extract spectral features and predict the onset and offset moments separately. The frame-level probabilities are inferred first, and then the note-level transcription results are obtained through post-processing. In this paper, a new AST framework called MusicYOLO is proposed, which obtains the note-level transcription results directly. The onset/offset detection is based on the object detection model YOLOX, and the pitch labeling is completed by a spectrogram peak search. Compared with previous methods, the MusicYOLO detects note objects rather than isolated onset/offset moments, thus greatly enhancing the transcription performance. On the sight-singing vocal dataset (SSVD) established in this paper, the MusicYOLO achieves an 84.60% transcription F1-score, which is the state-of-the-art method.
DOI: --
发表时间: 2016
期刊: --
影响因子: --
作者:
François Rigaud;Mathieu Radenen
通讯作者: François Rigaud;Mathieu Radenen
DOI: 10.1007/978-3-540-68585-2_34
发表时间: 2007
期刊: --
影响因子: --
作者:
A. Temko;C. Nadeu;Joan-Isaac Biel
通讯作者: A. Temko;C. Nadeu;Joan-Isaac Biel
DOI: 10.1109/icassp.2015.7178034
发表时间: 2015-04
期刊: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子: --
作者:
Yukara Ikemiya;Kazuyoshi Yoshii;Katsutoshi Itoyama
通讯作者: Yukara Ikemiya;Kazuyoshi Yoshii;Katsutoshi Itoyama
DOI: 10.1109/waspaa.2015.7336889
发表时间: 2015-11
期刊: 2015 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
影响因子: --
作者:
Huy Phan;M. Maass;L. Hertel;Radoslaw Mazur;A. Mertins
通讯作者: Huy Phan;M. Maass;L. Hertel;Radoslaw Mazur;A. Mertins
DOI: 10.1109/taslp.2016.2530401
发表时间: 2016-04-01
影响因子: 5.4
作者:
Huy Phan;Hertel, Lars;Mertins, Alfred
通讯作者: Mertins, Alfred