Inferring the Structure of a Tennis Game Using Audio Information

Inferring the Structure of a Tennis Game Using Audio Information
复制标题

DOI:
10.1109/tasl.2010.2103059
复制
发表时间:
2011-09
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Qiang Huang;S. Cox
Qiang Huang;S. Cox
中科院分区:
其他
文献类型:
--
作者:
Qiang Huang;S. Cox

文献摘要

被引文献

相似文献

我们描述了一个新的框架,用于推断的低层次结构的体育游戏(网球),只使用的信息上的音频记录的游戏。我们的目标是将比赛分割成一系列的点,这是描述网球比赛的自然单位。该框架是分层的,在最低级别包括音频事件的识别,然后是“匹配”(即,语义)事件,并且在最高级别,游戏点。不同的技术,适合于这些事件的每一个的特性被用来检测它们,这些技术耦合在一个概率框架。该技术包括高斯混合模型和分层语言模型来检测音频事件的序列,最大熵马尔可夫模型来推断“匹配”事件从这些音频事件和multigram来推断分割的匹配事件的序列到一个网球比赛中的点的序列。我们的结果是有希望的,给出了> 0.7的点的最终检测的F分数。
We describe a novel framework for inferring the low-level structure of a sports game (tennis) using only the information available on the audio track of a video recording of the game. Our goal is to segment the games into a sequence of points, the natural unit for describing a tennis match. The framework is hierarchical, consisting of, at the lowest level, identification of audio events, followed by “match” (i.e., semantic) events and at the highest level, game points. Different techniques that are appropriate to the characteristics of each of these events are used to detect them and these techniques are coupled in a probabilistic framework. The techniques consist of Gaussian mixture models and a hierarchical language model to detect sequences of audio events, a maximum entropy Markov model to infer “match” events from these audio events and multigrams to infer the segmentation of a sequence of match events into sequences of points in a a tennis game. Our results are promising, giving an F-score for the final detection of points of >; 0.7.