Do Acoustic Word Embeddings Capture Phonological Similarity? An Empirical Study

Do Acoustic Word Embeddings Capture Phonological Similarity? An Empirical Study
复制标题

DOI:
10.21437/interspeech.2021-678
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Badr M. Abdullah;Marius Mosbach;Iuliia Zaitova;B. Mobius;D. Klakow
Badr M. Abdullah;Marius Mosbach;Iuliia Zaitova;B. Mobius;D. Klakow
中科院分区:
其他
文献类型:
--
作者:
Badr M. Abdullah;Marius Mosbach;Iuliia Zaitova;B. Mobius;D. Klakow

文献摘要

被引文献

相似文献

深度神经网络的几种变体已被成功地用于构建参数模型,该模型将可变持续时间的口语词段投影到固定大小的矢量表示或声学单词嵌入(AWE)上。然而,目前还不清楚我们可以在多大程度上依赖新兴AWE空间中的距离作为词形相似性的估计。在本文中,我们问:声学嵌入空间中的距离是否与语音差异相关?为了回答这个问题,我们对具有不同神经结构和学习目标的AWE的监督方法的性能进行了实证研究。我们在两种语言(德语和捷克语)的受控环境中训练AWE模型,并评估嵌入在两个任务上的情况:单词辨别和语音相似性。我们的实验表明:(1)在最好的情况下,嵌入空间中的距离仅与语音距离适度相关;(2)提高单词辨别任务的成绩并不一定会产生更好地反映单词语音相似性的模型。我们的发现强调了重新思考目前对AWES的内在评估的必要性。
Several variants of deep neural networks have been successfully employed for building parametric models that project variable-duration spoken word segments onto fixed-size vector representations, or acoustic word embeddings (AWEs). However, it remains unclear to what degree we can rely on the distance in the emerging AWE space as an estimate of word-form similarity. In this paper, we ask: does the distance in the acoustic embedding space correlate with phonological dissimilarity? To answer this question, we empirically investigate the performance of supervised approaches for AWEs with different neural architectures and learning objectives. We train AWE models in controlled settings for two languages (German and Czech) and evaluate the embeddings on two tasks: word discrimination and phonological similarity. Our experiments show that (1) the distance in the embedding space in the best cases only moderately correlates with phonological distance, and (2) improving the performance on the word discrimination task does not necessarily yield models that better reflect word phonological similarity. Our findings highlight the necessity to rethink the current intrinsic evaluations for AWEs.