Transcribing Lead Sheet-Like Chord Progressions of Jazz Recordings

Transcribing Lead Sheet-Like Chord Progressions of Jazz Recordings
复制标题

转录爵士乐录音的主乐谱状和弦进行

DOI:
--
复制
发表时间:
2020
影响因子:
--
通讯作者:
P. Cuadra
P. Cuadra
中科院分区:
计算机科学4区
文献类型:
--
作者:
Gabriel Durán;P. Cuadra

文献摘要

被引文献

相似文献

绝大多数关于自动和弦转录的研究都是在数据库上开发和测试的,主要集中在流行和摇滚等流派上。然而,爵士乐强烈地基于即兴创作,并且和声的解释方式与许多其他流派不同,导致最先进的和弦转录系统表现不佳。本文提出了一个计算系统,转录爵士乐录音和弦,解决他们提出的具体挑战,并考虑其固有的音乐方面。采用原始音频和从用户手动获得的小调输入,系统可以联合转录和弦并检测录音的节拍,从而允许作为输出的铅板状渲染。分析分两部分进行。首先,使用动态时间扭曲基于其音乐内容对齐具有重复和弦进行(合唱)的所有片段。其次,对齐的片段被混合,卷积递归神经网络被用来同时检测节拍和转录和弦。这个自动和弦转录系统仅在爵士乐录音上训练和测试,并且比在非爵士乐特定的大型数据库上训练的其他系统实现了更好的性能。此外,它还结合了节拍检测和和弦转录任务,允许创建一个易于研究人员和音乐家解释的铅板状表示。
Abstract The vast majority of research on automatic chord transcription has been developed and tested on databases mainly focused on genres like pop and rock. Jazz is strongly based on improvisation, however, and the way harmony is interpreted is different from many other genres, causing state-of-the-art chord transcription systems to achieve poor performance. This article presents a computational system that transcribes chords from jazz recordings, addressing the specific challenges they present and considering their inherent musical aspects. Taking the raw audio and minor manually obtained inputs from the user, the system can jointly transcribe chords and detect the beat of a recording, allowing a lead sheet–like rendering as output. The analysis is implemented in two parts. First, all segments with a repeating chord progression (the chorus) are aligned based on their musical content using dynamic time warping. Second, the aligned segments are mixed and a convolutional recurrent neural network is used to simultaneously detect beats and transcribe chords. This automatic chord transcription system is trained and tested on jazz recordings only, and achieves better performance than other systems trained on larger databases that are not jazz specific. Additionally, it combines the beat-detection and chord transcription tasks, allowing the creation of a lead sheet–like representation that is easy to interpret by both researchers and musicians.