Talking Detection In Collaborative Learning Environments

Talking Detection In Collaborative Learning Environments
复制标题

DOI:
10.1007/978-3-030-89131-2_22
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Wenjing Shi;M. Pattichis;Sylvia Celedón-Pattichis;Carlos López Leiva
Wenjing Shi;M. Pattichis;Sylvia Celedón-Pattichis;Carlos López Leiva
中科院分区:
其他
文献类型:
--
作者:
Wenjing Shi;M. Pattichis;Sylvia Celedón-Pattichis;Carlos López Leiva

文献摘要

被引文献

相似文献

我们研究了协作学习视频中说话活动的检测问题。我们的方法使用头部检测和光流矢量的对数幅度的投影,以减少问题的小投影图像的简单分类,而不需要训练复杂的,3-D活动分类系统。然后,使用标准分类器的简单多数表决来容易地对小投影图像进行分类。对于说话检测,我们提出的方法显着优于单活动系统。我们的总体准确率为59%,而时间段网络(TSN)为42%,卷积3D(C3 D)为45%。此外,我们的方法能够检测来自多个扬声器的多个说话实例,同时也检测扬声器本身。
We study the problem of detecting talking activities in collaborative learning videos. Our approach uses head detection and projections of the log-magnitude of optical flow vectors to reduce the problem to a simple classification of small projection images without the need for training complex, 3-D activity classification systems. The small projection images are then easily classified using a simple majority vote of standard classifiers. For talking detection, our proposed approach is shown to significantly outperform single activity systems. We have an overall accuracy of 59% compared to 42% for Temporal Segment Network (TSN) and 45% for Convolutional 3D (C3D). In addition, our method is able to detect multiple talking instances from multiple speakers, while also detecting the speakers themselves.