Particle-filter based audio-visual beat-tracking for music robot ensemble with human guitarist

Particle-filter based audio-visual beat-tracking for music robot ensemble with human guitarist
复制标题

DOI:
10.1109/iros.2011.6094773
复制
发表时间:
2011-12
期刊:
2011 IEEE/RSJ International Conference on Intelligent Robots and Systems
影响因子:
--
通讯作者:
Tatsuhiko Itohara;Takuma Otsuka;Takeshi Mizumoto;T. Ogata;HIroshi G. Okuno
Tatsuhiko Itohara;Takuma Otsuka;Takeshi Mizumoto;T. Ogata;HIroshi G. Okuno
中科院分区:
其他
文献类型:
--
作者:
Tatsuhiko Itohara;Takuma Otsuka;Takeshi Mizumoto;T. Ogata;HIroshi G. Okuno

文献摘要

相似文献

本文提出了一种视听节拍跟踪方法合奏机器人与人类吉他手。节拍跟踪,或音乐的克里思和节拍时间的估计,对于高质量的音乐合奏表演至关重要。由于人类在后拍和切分中以外拍演奏吉他,因此人类吉他演奏的节拍跟踪的主要问题是双重的:克里思变化和变化的音符长度。大多数传统的方法没有解决人类的吉他演奏。因此,它们缺乏对任何一个问题的适应。为了同时解决这些问题,我们的方法不仅使用音频,而且使用视觉特征。我们提取音频特征的频谱-时间模式匹配(STPM)和视觉特征的光流,均值漂移和Hough变换。我们的节拍跟踪使用粒子滤波器估计克里思和节拍时间;吉他声音的声学特征和手臂运动的视觉特征都表示为粒子。实验结果证实,我们的综合视听方法对克里思变化和不同的音符长度是鲁棒的。此外,他们还表明,估计的收敛速度只依赖于粒子的数量很少。当粒子数为200时,实时因子为0.88,这表明该方法是实时的。
This paper presents an audio-visual beat-tracking method for ensemble robots with a human guitarist. Beat-tracking, or estimation of tempo and beat times of music, is critical to the high quality of musical ensemble performance. Since a human plays the guitar in out-beat in back beat and syncopation, the main problems of beat-tracking of a human's guitar playing are twofold: tempo changes and varying note lengths. Most conventional methods have not addressed human's guitar playing. Therefore, they lack the adaptation of either of the problems. To solve the problems simultaneously, our method uses not only audio but visual features. We extract audio features with Spectro-Temporal Pattern Matching (STPM) and visual features with optical flow, mean shift and Hough transform. Our beat-tracking estimates tempo and beat time using a particle filter; both acoustic feature of guitar sounds and visual features of arm motions are represented as particles. The particle is determined based on prior distribution of audio and visual features, respectively Experimental results confirm that our integrated audio-visual approach is robust against tempo changes and varying note lengths. In addition, they also show that estimation convergence rate depends only a little on the number of particles. The real-time factor is 0.88 when the number of particles is 200, and this shows out method works in real-time.