DEEP-HEAR: A Multimodal Subtitle Positioning System Dedicated to Deaf and Hearing-Impaired People

DEEP-HEAR: A Multimodal Subtitle Positioning System Dedicated to Deaf and Hearing-Impaired People
复制标题

DOI:
10.1109/access.2019.2925806
复制
发表时间:
2019
期刊:
影响因子:
3.9
通讯作者:
Ruxandra Tapu;B. Mocanu;T. Zaharia
Ruxandra Tapu;B. Mocanu;T. Zaharia
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ruxandra Tapu;B. Mocanu;T. Zaharia

文献摘要

被引文献

相似文献

本文介绍了一种多模态动态字幕定位系统DEEP-HEAR框架,该系统旨在提高聋人和听障人士(HIP)对多媒体文档的可及性。该系统利用计算机视觉算法和深度卷积神经网络,专门设计和调整,以检测和识别主动说话者的身份。本文的主要贡献在于:提出了一种新的方法来识别视频流中存在的各种字符。一种基于人脸轨迹和视觉一致性将视频序列划分为语义单元的视频时间分割算法。最后,我们的方法的核心是一种基于文本、音频和视频流的多模态信息融合的新型主动说话人识别方法。在30多个视频的大规模数据集上进行的实验结果验证了所提出的方法,平均准确率和识别率优于90%。此外,该方法对重要的物体/相机运动和面部姿势变化具有鲁棒性,与最先进的技术相比,精确度和召回率提高了8%以上。通过对动态字幕定位系统的主观评价,验证了该方法的有效性。
In this paper, we introduce the DEEP-HEAR framework, a multimodal dynamic subtitle positioning system designed to increase the accessibility of deaf and hearing impaired people (HIP) to multimedia documents. The proposed system exploits both computer vision algorithms and deep convolutional neural networks specifically designed and tuned in order to detect and recognize the identity of the active speaker. The main contributions of the paper concern: a novel method dedicated to recognizing various characters existent in the video stream. A video temporal segmentation algorithm that divides the video sequence into semantic units, based on face tracks and visual consistency. Finally, the core of our approach concerns a novel active speaker recognition method relying on the multimodal information fusion from the text, audio, and video streams. The experimental results carried out on a large scale dataset of more than 30 videos, validate the proposed methodology with average accuracy and recognition rates superior to 90%. Moreover, the method shows robustness to important object/camera motion and face pose variation, yielding gains of more than 8% in precision and recall rates when compared with state-of-the-art techniques. The subjective evaluation of the proposed dynamic subtitle positioning system demonstrates the effectiveness of our approach.