Generalized framework for summarization of fixed-camera lecture videos by detecting and binarizing handwritten content

Generalized framework for summarization of fixed-camera lecture videos by detecting and binarizing handwritten content
复制标题

DOI:
10.1007/s10032-019-00327-y
复制
发表时间:
2019-06
期刊:
International Journal on Document Analysis and Recognition (IJDAR)
影响因子:
--
通讯作者:
B. Kota;Kenny Davila;Alexander Stone;S. Setlur;V. Govindaraju
B. Kota;Kenny Davila;Alexander Stone;S. Setlur;V. Govindaraju
中科院分区:
其他
文献类型:
--
作者:
B. Kota;Kenny Davila;Alexander Stone;S. Setlur;V. Govindaraju

文献摘要

被引文献

相似文献

我们提出了一个框架,提取和二进制化的手写内容的讲座视频。提取的内容可能会被用于索引视频集合,从而在讲座视频中提供基于内容的搜索和导航,帮助世界各地的学生和教育工作者。深度学习管道用于检测手写文本、公式和草图,然后将提取的内容二进制化。我们利用我们的二值化检测的时空结构来计算跨所有视频帧的内容的关联性信息。该信息稍后用于分割视频。实验进行比较我们的框架中的关键组件的性能隔离,以及对整体性能的影响,相对于现有的方法。我们评估我们的框架上公开提供的数学讲座视频数据集获得的f-措施的二进制连接组件。框架代码(包括训练的权重)和总结将被发布。
We propose a framework to extract and binarize handwritten content in lecture videos. The extracted content could potentially be used to index video collections powering content-based search and navigation within lecture videos helping students and educators across the world. A deep learning pipeline is used to detect handwritten text, formulae and sketches and then binarize the extracted content. We exploit the spatio-temporal structure of our binarized detections to compute associativity information of content across all video frames. This information is later used to segment the video. Experiments are conducted to compare the performance of key components of our framework in isolation, as well as the impact on overall performance, with respect to existing methods. We evaluate our framework on the publicly available AccessMath lecture video dataset obtaining anf-measure offor binary connected components. Code for the framework (including trained weights) and summarization will be released.