SpecTextor: End-to-End Attention-based Mechanism for Dense Text Generation in Sports Journalism

SpecTextor: End-to-End Attention-based Mechanism for Dense Text Generation in Sports Journalism
复制标题

DOI:
10.1109/smartcomp55677.2022.00081
复制
发表时间:
2022-06
期刊:
2022 IEEE International Conference on Smart Computing (SMARTCOMP)
影响因子:
--
通讯作者:
Indrajeet Ghosh;Matthew Ivler;S. R. Ramamurthy;Nirmalya Roy
Indrajeet Ghosh;Matthew Ivler;S. R. Ramamurthy;Nirmalya Roy
中科院分区:
其他
文献类型:
--
作者:
Indrajeet Ghosh;Matthew Ivler;S. R. Ramamurthy;Nirmalya Roy

文献摘要

相似文献

智能导航系统可以帮助设计下一代人机交互应用程序。密集文本描述是系统学习每个视频帧的语义知识和视觉特征并将其映射以描述视频最相关的主题和事件的研究领域之一。在本文中,我们认为未经修剪的体育视频作为我们的案例研究。在体育领域产生密集的描述,以补充新闻作品,而不依赖于评论员和专家,需要更多的调查。出于这一动机,我们提出了一个端到端的自动文本生成器,SpecTextor,学习的语义特征,从体育比赛的未修剪的视频,并生成相关的描述性文本。所提出的方法认为,视频作为一个序列的帧,并按顺序生成的话。在将视频分割成帧后,我们使用预训练的VGG-16模型进行特征提取和视频帧编码。通过这些编码帧,我们构建了一个基于长短期记忆(LSTM)的注意力解码器管道,该管道利用软注意力机制将语义特征与相关文本描述进行映射,以生成游戏的解释。由于开发游戏的全面描述需要在一组密集的时间戳字幕上进行训练,因此我们利用了两个可用的公共数据集:ActivityNet字幕和Microsoft视频描述。此外,我们使用了两种不同的解码算法:波束搜索和贪婪搜索,并计算了两个评估指标:BLEU和METEOR分数。
Language-guided smart systems can help to design next-generation human-machine interactive applications. The dense text description is one of the research areas where systems learn the semantic knowledge and visual features of each video frame and map them to describe the video's most relevant subjects and events. In this paper, we consider untrimmed sports videos as our case study. Generating dense descriptions in the sports domain to supplement journalistic works without relying on commentators and experts requires more investigation. Motivated by this, we propose an end-to-end automated text-generator, SpecTextor, that learns the semantic features from untrimmed videos of sports games and generates associated descriptive texts. The proposed approach considers the video as a sequence of frames and sequentially generates words. After splitting videos into frames, we use a pre-trained VGG-16 model for feature extraction and encoding the video frames. With these encoded frames, we posit a Long Short-Term Memory (LSTM) based attention-decoder pipeline that leverages soft-attention mechanism to map the semantic features with relevant textual descriptions to generate the explanation of the game. Because developing a comprehensive description of the game warrants training on a set of dense time-stamped captions, we leverage two available public datasets: ActivityNet Captions and Microsoft Video Description. In addition, we utilized two different decoding algorithms: beam search and greedy search and computed two evaluation metrics: BLEU and METEOR scores.