Bridging the Gap: Using Deep Acoustic Representations to Learn Grounded Language from Percepts and Raw Speech

Bridging the Gap: Using Deep Acoustic Representations to Learn Grounded Language from Percepts and Raw Speech
复制标题

DOI:
10.1609/aaai.v36i10.21335
复制
发表时间:
2021-12
期刊:
--
影响因子:
--
通讯作者:
Gaoussou Youssouf Kebe;Luke E. Richards;Edward Raff;Francis Ferraro;Cynthia Matuszek
Gaoussou Youssouf Kebe;Luke E. Richards;Edward Raff;Francis Ferraro;Cynthia Matuszek
中科院分区:
其他
文献类型:
--
作者:
Gaoussou Youssouf Kebe;Luke E. Richards;Edward Raff;Francis Ferraro;Cynthia Matuszek

文献摘要

相似文献

学习理解扎根的语言是一个关键的研究领域,它将自然语言与感知联系起来。扎根语言习得方面的研究主要集中在文本输入上。在这项工作中,我们论证了在配对的视觉感知和原始语音输入上进行扎根语言习得的可行性。这将允许人与机器人进行互动,从最终用户那里学习关于新任务和环境的语言,减少对文本输入的依赖,并潜在地缓解广泛可用的语音识别系统中存在的人口统计学偏见的影响。我们利用最近在自我监督语音表示模型方面的工作,并表明学习的语音表示可以使语言接地系统对特定群体更具包容性,同时保持甚至提高总体性能。
Learning to understand grounded language, which connects natural language to percepts, is a critical research area. Prior work in grounded language acquisition has focused primarily on textual inputs. In this work, we demonstrate the feasibility of performing grounded language acquisition on paired visual percepts and raw speech inputs. This will allow human-robot interactions in which language about novel tasks and environments is learned from end-users, reducing dependence on textual inputs and potentially mitigating the effects of demographic bias found in widely available speech recognition systems. We leverage recent work in self-supervised speech representation models and show that learned representations of speech can make language grounding systems more inclusive towards specific groups while maintaining or even increasing general performance.