Automatic Identification for Singing Style based on Sung Melodic Contour Characterized in Phase Plane

Automatic Identification for Singing Style based on Sung Melodic Contour Characterized in Phase Plane
复制标题

基于相平面歌唱旋律轮廓的演唱风格自动识别

DOI:
--
复制
发表时间:
2009
期刊:
International Society for Music Information Retrieval Conference
影响因子:
--
通讯作者:
K. Takeda
K. Takeda
中科院分区:
--
文献类型:
--
作者:
Tatsuya Kako;Yasunori Ohishi;H. Kameoka;K. Kashino;K. Takeda

文献摘要

被引文献

相似文献

提出了一种歌唱风格的随机表示方法.旋律轮廓的动态性,即,基频(F0)序列是歌唱风格的主要线索,因为它可以表征颤音等典型的表现形式。相平面中的F0信号轨迹被用作基本表示。通过将高斯混合模型拟合到相平面上观测到的F0轨迹上,由一组GMM参数获得参数表示。实验结果表明,该方法的有效性得到了证实,单类判别准确率达到94.1%。这些研究都试图以旋律轮廓的局部动态作为演唱的线索,没有提出系统的方法来表征演唱风格。在(14,17 -19)中报道了典型表现的滞后系统模型;然而,没有讨论歌唱风格的变化。在本文中,我们提出了一个随机相平面作为歌唱风格的图形表示,并显示其有效性歌唱风格的歧视。这种表征歌唱风格的表示的一个优点是,由于既不需要用于像颤音的重复的显式检测函数,也不需要对目标音符的估计,所以它对演唱的旋律是鲁棒的。在先前的论文(20)中,我们将相平面中F0轮廓的这种图形表示应用于Hamming查询系统,并中和了F0序列的局部动态,使得查询仅利用音乐信息。与此相反,由于音乐信息与演唱风格是一种二元关系,因此本研究采用F0序列的局部动力学来模拟演唱风格,而忽略了音乐信息。在本文中,我们还评估了所提出的表示,通过一个歌手类的歧视实验中,我们表明,我们提出的模型可以提取的动态特性的演唱旋律共享的一组歌手。在下一节中,我们提出了随机相平面(SPP)作为旋律轮廓的随机表示,并展示了如何通过所提出的SPP对歌唱表示进行建模。在第3节中,我们通过歌手类判别实验证明了我们所提出的方法的有效性。第四节讨论了所得结果并对本文进行了总结。
A stochastic representation of singing styles is pro- posed. The dynamic property of melodic contour, i.e., fun- damental frequency (F0) sequence, is assumed to be the main cue for singing styles because it can characterize such typical ornamentations as vibrato . F0 signal trajectories in the phase plane are used as the basic representation. By fitting Gaussian mixture models to the observed F0 trajec- tories in the phase plane, a parametric representation is ob- tained by a set of GMM parameters. The effectiveness of our proposed method is confirmed through experimental evaluation where 94.1% accuracy for singer-class discrim- ination was obtained. these studies try to use the local dynamics of melodic con- tour as a cue for ornamentation, no systematic method has been proposed for characterizing singing styles. A lag system model for typical ornamentations was reported in (14,17-19); however, variation of singing styles was not discussed. In this paper, we propose a stochastic phase plane as a graphical representation of singing styles and show its effectiveness for singing style discrimination. One merit of this representation to characterize singing style is that since neither an explicit detection function for ornamen- tation like vibrato nor estimation of the target note is re- quired, it is robust to sung melodies. In a previous paper (20), we applied this graphical rep- resentation of the F0 contour in the phase plane to a query- by-hamming system and neutralized the local dynamics of the F0 sequence so that only musical information was uti- lized for the query. In contrast, in this study, we use the local dynamics of the F0 sequence for modeling singing styles and disregard the musical information because mu- sical information and singing style are in a dual relation. In this paper, we also evaluate the proposed represen- tation through a singer-class discrimination experiment in which we show that our proposed model can extract the dynamic properties of sung melodies shared by a group of singers. In the next section, we propose stochastic phase plane (SPP) as a stochastic representation of the melodic contour and show how singing ornamentations are modeled by the proposed SPP. In Section 3, we experimentally show the effectiveness of our proposed method through singer class discrimination experiments. Section 4 discusses the ob- tained results and concludes this paper.