Bilingual Speech Recognition by Estimating Speaker Geometry from Video Data

Bilingual Speech Recognition by Estimating Speaker Geometry from Video Data
复制标题

通过从视频数据估计说话者几何形状进行双语语音识别

DOI:
10.1007/978-3-030-89128-2_8
复制
发表时间:
2021
期刊:
Computer Analysis of Images and Patterns. CAIP 2021. Lecture Notes in Computer Science.
影响因子:
--
通讯作者:
LópezLeiva, C.
LópezLeiva, C.
中科院分区:
--
文献类型:
--
作者:
Tapia, L.;Gomez, A.;Esparza, M.;Jatla, V.;Pattichis, M.;Celedón-Pattichis, S.;LópezLeiva, C.

文献摘要

参考文献

被引文献

相似文献

语音识别是非常具有挑战性的学生学习环境,其特点是显着的串扰和背景噪声。为了解决这个问题,我们提出了一个双语语音识别系统,使用交互式视频分析系统来估计逼真的音频模拟的3D扬声器的几何形状。我们演示了使用我们的系统生成一个复杂的音频数据集,其中包含显着的串扰和背景噪声,近似于现实生活中的教室录音。然后,我们测试我们提出的系统与现实生活中的录音。在扬声器从麦克风的距离方面,我们的交互式视频分析系统获得了更好的平均错误率为10.83%相比,基线方法的33.12%。我们提出的系统在同一数据集上的准确率为27.92%,比Google Speech-to-text高1.5%。在9个重要关键词方面,我们的方法的平均灵敏度为38%,而Google Speech-to-text的平均灵敏度为24%,而两种方法的平均特异度分别为90%和92%。另一方面,两种方法的特异性仍然很高(90%至92%)。
Speech recognition is very challenging in student learning environments that are characterized by significant cross-talk and background noise. To address this problem, we present a bilingual speech recognition system that uses an interactive video analysis system to estimate the 3D speaker geometry for realistic audio simulations. We demonstrate the use of our system in generating a complex audio dataset that contains significant cross-talk and background noise that approximate real-life classroom recordings. We then test our proposed system with real-life recordings.In terms of the distance of the speakers from the microphone, our interactive video analysis system obtained a better average error rate of 10.83% compared to 33.12% for a baseline approach. Our proposed system gave an accuracy of 27.92% that is 1.5% better than Google Speech-to-text on the same dataset. In terms of 9 important keywords, our approach gave an average sensitivity of 38% compared to 24% for Google Speech-to-text, while both methods maintained high average specificity of 90% and 92%.On average, sensitivity improved from 24% to 38% for our proposed approach. On the other hand, specificity remained high for both methods (90% to 92%).
协作学习视频中的动态小组互动
DOI: 10.1109/acssc.2018.8645132
发表时间: 2018
期刊: and Computers
影响因子: --
作者:
Shi, Wenjing;Pattichis, Marios S.;Celedon-Pattichis, Sylvia;LopezLeiva, Carlos
通讯作者: LopezLeiva, Carlos
DOI: 10.1007/978-3-030-89131-2_22
发表时间: 2021-10
期刊: --
影响因子: --
作者:
Wenjing Shi;M. Pattichis;Sylvia Celedón-Pattichis;Carlos López Leiva
通讯作者: Wenjing Shi;M. Pattichis;Sylvia Celedón-Pattichis;Carlos López Leiva
DOI: --
发表时间: 2021
期刊: and Computers.
影响因子: --
作者:
Jatla, V.;Teeparthi, S.;Pattichis, M.S.;Celedon-Pattichis, S.;LopezLeiva, C.
通讯作者: LopezLeiva, C.
DOI: 10.1109/ssiai.2018.8470331
发表时间: 2018
期刊: 2018 IEEE Southwest Symposium on Image Analysis and Interpretation (SSIAI)
影响因子: --
作者:
A. Jacoby;M. Pattichis;Sylvia Celedón;Carlos A. LópezLeiva
通讯作者: Carlos A. LópezLeiva
DOI: --
发表时间: 2013
期刊:
影响因子: --
作者:
Sylvia Celedón;Carlos A. LópezLeiva;M. Pattichis;D. Llamocca
通讯作者: D. Llamocca