Multimodal framework based on audio-visual features for summarisation of cricket videos

Multimodal framework based on audio-visual features for summarisation of cricket videos
复制标题

DOI:
10.1049/iet-ipr.2018.5589
复制
发表时间:
2019-03-28
影响因子:
2.3
通讯作者:
Adnan, Syed
Adnan, Syed
中科院分区:
计算机科学4区
文献类型:
--
作者:
Javed, Ali;Irtaza, Aun;Adnan, Syed

文献摘要

被引文献

相似文献

由于世界各地的大量观众,体育广播公司在网络空间上产生了大量的视频内容。分析和使用这个巨大的存储库促使广播公司应用视频摘要来从整个视频中提取令人兴奋的片段,以捕捉用户的兴趣并获得存储和传输的好处。因此,在这项研究中的关键事件检测和总结的基础上的视听功能的板球视频的自动方法。声学局部二进制模式特征用于捕获音频流中的兴奋水平,该兴奋水平用于训练二进制支持向量机(SVM)分类器。训练的SVM分类器用于将音频帧标记为激励或非激励帧。激励音频帧用于选择候选关键视频帧。一个基于决策树的分类器被训练来检测输入板球视频中的关键事件,然后将其用于视频摘要。所提出的框架的性能进行了评估属于不同的比赛和广播公司的板球视频的不同数据集。实验结果表明,该方法的平均准确率达到95.5%,这表明其有效性。
Sports broadcasters generate an enormous amount of video content on the cyberspace due to massive viewership all over the world. Analysis and consumption of this huge repository urges the broadcasters to apply video summarisation to extract the exciting segments from the entire video to capture user's interest and reap the storage and transmission benefits. Therefore, in this study an automatic method for key-events detection and summarisation based on audio-visual features is presented for cricket videos. Acoustic local binary pattern features are used to capture excitement level in the audio stream, which is used to train a binary support vector machine (SVM) classifier. Trained SVM classifier is used to label audio frame as an excited or non-excited frame. Excited audio frames are used to select candidate key-video frames. A decision tree-based classifier is trained to detect key-events in the input cricket videos that are then used for video summarisation. Performance of the proposed framework has been evaluated on a diverse dataset of cricket videos belonging to different tournaments and broadcasters. Experimental results indicate that the proposed method achieves an average accuracy of 95.5%, which signifies its effectiveness.