How Noisy is Too Noisy? The Impact of Data Noise on Multimodal Recognition of Confusion and Conflict During Collaborative Learning

How Noisy is Too Noisy? The Impact of Data Noise on Multimodal Recognition of Confusion and Conflict During Collaborative Learning
复制标题

DOI:
10.1145/3577190.3614127
复制
发表时间:
2023-10
期刊:
Proceedings of the 25th International Conference on Multimodal Interaction
影响因子:
--
通讯作者:
Yingbo Ma;Mehmet Celepkolu;K. Boyer;Collin Lynch;E. Wiebe;Maya Israel
Yingbo Ma;Mehmet Celepkolu;K. Boyer;Collin Lynch;E. Wiebe;Maya Israel
中科院分区:
其他
文献类型:
--
作者:
Yingbo Ma;Mehmet Celepkolu;K. Boyer;Collin Lynch;E. Wiebe;Maya Israel

文献摘要

相似文献

支持协作学习的智能系统依赖于实时行为数据,包括语言,音频和视频。多模式数据的实用性是我们如何在本文中构建可靠的多模型的一个开放式问题。 25个小学学习者二元组的协作编程会议期间的冲突时刻。准确性。结果表明,当WER超过20%时,模型检测语言方式的混乱和冲突的准确性急剧下降到0.73。同样,在音频方式中,当SNR降至5 dB以下时,模型的精度从0.79急剧下降到0.61。此外,学习者的脸得到了成功,我们训练了多种模型非模式的数据,最终导致了确认混乱和冲突的准确性。
Intelligent systems to support collaborative learning rely on real-time behavioral data, including language, audio, and video. However, noisy data, such as word errors in speech recognition, audio static or background noise, and facial mistracking in video, often limit the utility of multimodal data. It is an open question of how we can build reliable multimodal models in the face of substantial data noise. In this paper, we investigate the impact of data noise on the recognition of confusion and conflict moments during collaborative programming sessions by 25 dyads of elementary school learners. We measure language errors with word error rate (WER), audio noise with speech-to-noise ratio (SNR), and video errors with frame-by-frame facial tracking accuracy. The results showed that the model’s accuracy for detecting confusion and conflict in the language modality decreased drastically from 0.84 to 0.73 when the WER exceeded 20%. Similarly, in the audio modality, the model’s accuracy decreased sharply from 0.79 to 0.61 when the SNR dropped below 5 dB. Conversely, the model’s accuracy remained relatively constant in the video modality at a comparable level (> 0.70) so long as at least one learner’s face was successfully tracked. Moreover, we trained several multimodal models and found that integrating multimodal data could effectively offset the negative effect of noise in unimodal data, ultimately leading to improved accuracy in recognizing confusion and conflict. These findings have practical implications for the future deployment of intelligent systems that support collaborative learning in actual classroom settings.