Automatic audiovisual synchronisation for ultrasound tongue imaging

Automatic audiovisual synchronisation for ultrasound tongue imaging
复制标题

DOI:
10.1016/j.specom.2021.05.008
复制
发表时间:
2021-06-12
影响因子:
3.2
通讯作者:
Renals, Steve
Renals, Steve
中科院分区:
计算机科学3区
文献类型:
--
作者:
Eshky, Aciel;Cleland, Joanne;Renals, Steve

文献摘要

被引文献

相似文献

超声舌成像用于在语音产生期间可视化口内发音器官。它被用于一系列应用,包括言语和语言治疗以及语音学研究。超声和语音音频是同时记录的,为了正确使用这些数据,这两种模式应该正确同步。在记录时使用专用硬件实现同步,但这种方法在实践中可能失败,导致数据的可用性有限。在本文中,我们解决了数据收集后自动同步超声和音频的问题。我们首先调查专家超声用户同步错误的容忍度,以找到错误检测的阈值。我们使用这些阈值来定义准确性评分边界,以评估我们的系统。然后,我们描述了我们的自动同步方法,该方法由自监督神经网络驱动,利用两个信号之间的相关性来同步它们。我们在来自不同扬声器特征、不同设备和不同录音环境的多个域的数据上训练我们的模型,并在保持的域内数据上实现了>92.4%的准确率。最后,我们介绍了一种新的资源,裂缝数据集,我们收集了一个新的临床亚组,硬件同步证明是不可靠的。我们将我们的模型应用于这个域外数据,并与专家用户主观地评估其性能。结果表明,用户更喜欢我们的模型的输出超过原来的硬件输出的79.3%的时间。我们的研究结果表明,我们的方法的力量和能力,从新的领域推广到数据。
Ultrasound tongue imaging is used to visualise the intra-oral articulators during speech production. It is utilised in a range of applications, including speech and language therapy and phonetics research. Ultrasound and speech audio are recorded simultaneously, and in order to correctly use this data, the two modalities should be correctly synchronised. Synchronisation is achieved using specialised hardware at recording time, but this approach can fail in practice resulting in data of limited usability. In this paper, we address the problem of automatically synchronising ultrasound and audio after data collection. We first investigate the tolerance of expert ultrasound users to synchronisation errors in order to find the thresholds for error detection. We use these thresholds to define accuracy scoring boundaries for evaluating our system. We then describe our approach for automatic synchronisation, which is driven by a self-supervised neural network, exploiting the correlation between the two signals to synchronise them. We train our model on data from multiple domains with different speaker characteristics, different equipment, and different recording environments, and achieve an accuracy >92.4% on held-out in-domain data. Finally, we introduce a novel resource, the Cleft dataset, which we gathered with a new clinical subgroup and for which hardware synchronisation proved unreliable. We apply our model to this out-of-domain data, and evaluate its performance subjectively with expert users. Results show that users prefer our model's output over the original hardware output 79.3% of the time. Our results demonstrate the strength of our approach and its ability to generalise to data from new domains.