Hierarchical language modeling for audio events detection in a sports game

Hierarchical language modeling for audio events detection in a sports game
复制标题

DOI:
10.1109/icassp.2010.5495935
复制
发表时间:
2010-03
期刊:
2010 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
Qiang Huang;S. Cox
Qiang Huang;S. Cox
中科院分区:
其他
文献类型:
--
作者:
Qiang Huang;S. Cox

文献摘要

相似文献

我们从体育比赛的录音中研究“事件”的自动标签。我们描述了一种利用语言模型层次结构的技术,该语言模型是声音观察的低级模型和游戏中发生的音频事件的高级模型:这些模型使用最大熵方法集成。我们的音频事件模型还利用了持续时间和语音信息以及频谱内容,并且我们表明,使用这些特征可以进一步区分事件。不同网球比赛的结果表明,使用这些技术比使用不使用帧和事件之间的依赖关系建模或以持续时间和声音形式的额外信息的方法要好。
We investigate the automatic labelling of “events” from an audio recording of a sports game. We describe a technique that utilises a hierarchy of language models, which are a low-level model of acoustic observations and a high-level model of audio events that occur during a game: these models are integrated using a maximum entropy approach. Our models of the audio events also utilise duration and voicing information as well as spectral content, and we show that further discrimination between events is possible using these features. Results on different tennis games show that the use of these techniques is better than using an approach that does not use modelling of dependencies between frames and events or extra information in the form of duration and voicing.