A Framework for Multi-f0 Modeling in SATB Choir Recordings

A Framework for Multi-f0 Modeling in SATB Choir Recordings
复制标题

SATB 合唱团录音中的多 f0 建模框架

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
Pritish Chandna
Pritish Chandna
中科院分区:
--
文献类型:
--
作者:
Helena Cuesta;E. Gómez;Pritish Chandna

文献摘要

被引文献

相似文献

基频(f0)建模是一个重要的,但相对未开发的合唱团演唱方面。无论是个人演唱还是合唱团演唱,演唱的性能评估以及听觉分析通常都依赖于提取歌声的f0轮廓。然而,由于大量的歌手,在类似的频率范围内唱歌,从合唱团录音中提取精确的个人音高轮廓是一项具有挑战性的任务。在本文中,我们解决这一问题,并制定了一个方法来模拟音高轮廓的SATB合唱团录音。一个典型的SATB合唱团由四个部分组成,每个部分涵盖不同的音高范围,通常每个部分有多个歌手。我们首先评估一些国家的最先进的多f0估计系统的特定情况下,合唱团与一个单一的歌手每部分,并观察到,个别歌手的音高可以估计到一个相对较高的准确度。然而,我们观察到,每个合唱团部分(即齐声演唱)的多名歌手的情况更具挑战性。在这项工作中,我们提出了一种方法,该方法基于结合基于深度学习的多f0估计方法,然后是一组传统的DSP技术来建模f0及其分散,而不是每个合唱团部分的单个f0轨迹。我们提出并讨论了我们的意见和测试我们的框架与不同的歌手配置。
Fundamental frequency (f0) modeling is an important but relatively unexplored aspect of choir singing. Performance evaluation as well as auditory analysis of singing, whether individually or in a choir, often depend on extracting f0 contours for the singing voice. However, due to the large number of singers, singing at a similar frequency range, extracting the exact individual pitch contours from choir recordings is a challenging task. In this paper, we address this task and develop a methodology for modeling pitch contours of SATB choir recordings. A typical SATB choir consists of four parts, each covering a distinct range of pitches and often with multiple singers each. We first evaluate some state-of-the-art multi-f0 estimation systems for the particular case of choirs with a single singer per part, and observe that the pitch of individual singers can be estimated to a relatively high degree of accuracy. We observe, however, that the scenario of multiple singers for each choir part (i.e. unison singing) is far more challenging. In this work we propose a methodology based on combining a multi-f0 estimation methodology based on deep learning followed by a set of traditional DSP techniques to model f0 and its dispersion instead of a single f0 trajectory for each choir part. We present and discuss our observations and test our framework with different singer configurations.