Beyond the Edge: Markerless Pose Estimation of Speech Articulators from Ultrasound and Camera Images Using DeepLabCut.

Beyond the Edge: Markerless Pose Estimation of Speech Articulators from Ultrasound and Camera Images Using DeepLabCut.
复制标题

DOI:
10.3390/s22031133
复制
发表时间:
2022-02-02
期刊:
Sensors (Basel, Switzerland)
影响因子:
--
通讯作者:
Balch-Tomes J
Balch-Tomes J
中科院分区:
其他
文献类型:
--
作者:
Wrench A;Balch-Tomes J

文献摘要

参考文献

相似文献

语音发音器官图像的自动特征提取目前是通过检测边缘来实现的。在这里,我们研究了使用姿势估计深度神经网络和迁移学习来执行语音发音器关键点的无标记估计,仅使用几百个手动标记的图像作为训练输入。舌、颌和舌骨的正中矢状超声图像和嘴唇的相机图像用关键点手动标记,使用DeepLabCut进行训练,并在看不见的扬声器和系统上进行评价。从估计的和手工标记的关键点内插的舌表面轮廓产生0.93,s. d.的平均距离和(MSD)。0.46 mm,s.d. 0.39 mm,对于两个人标记者,和2.3,s. d. 1.5 mm,以获得最佳性能的边缘检测算法。一组同时进行的电磁关节造影(EMA)和超声记录的试验表明,三个物理传感器位置和相应的估计关键点之间存在部分相关性,需要进一步研究。从摄像机视频中估计唇口的准确性很高,平均MSD为0.70,s.d.。0.56 mm,s.d. 0.48 mm用于两个人类贴标机。DeepLabCut被发现是一种快速,准确和全自动的方法,可以为舌头,舌骨,下巴和嘴唇提供独特的运动学数据。
Automatic feature extraction from images of speech articulators is currently achieved by detecting edges. Here, we investigate the use of pose estimation deep neural nets with transfer learning to perform markerless estimation of speech articulator keypoints using only a few hundred hand-labelled images as training input. Midsagittal ultrasound images of the tongue, jaw, and hyoid and camera images of the lips were hand-labelled with keypoints, trained using DeepLabCut and evaluated on unseen speakers and systems. Tongue surface contours interpolated from estimated and hand-labelled keypoints produced an average mean sum of distances (MSD) of 0.93, s.d. 0.46 mm, compared with 0.96, s.d. 0.39 mm, for two human labellers, and 2.3, s.d. 1.5 mm, for the best performing edge detection algorithm. A pilot set of simultaneous electromagnetic articulography (EMA) and ultrasound recordings demonstrated partial correlation among three physical sensor positions and the corresponding estimated keypoints and requires further investigation. The accuracy of the estimating lip aperture from a camera video was high, with a mean MSD of 0.70, s.d. 0.56 mm compared with 0.57, s.d. 0.48 mm for two human labellers. DeepLabCut was found to be a fast, accurate and fully automatic method of providing unique kinematic data for tongue, hyoid, jaw, and lips.
DOI: 10.1038/s41593-018-0209-y
发表时间: 2018-09-01
影响因子: 25
作者:
Mathis, Alexander;Mamidanna, Pranav;Bethge, Matthias
通讯作者: Bethge, Matthias
DOI: 10.1016/j.conb.2019.10.008
发表时间: 2020-02-01
影响因子: 5.7
作者:
Mathis, Mackenzie Weygandt;Mathis, Alexander
通讯作者: Mathis, Alexander
DOI: 10.1007/bf00133570
发表时间: 1987-01-01
影响因子: 19.5
作者:
KASS, M;WITKIN, A;TERZOPOULOS, D
通讯作者: TERZOPOULOS, D
DOI: 10.1038/s41596-019-0176-0
发表时间: 2019-07-01
期刊: NATURE PROTOCOLS
影响因子: 14.8
作者:
Nath, Tanmay;Mathis, Alexander;Mathis, Mackenzie Weygandt
通讯作者: Mathis, Mackenzie Weygandt
DOI: 10.7554/elife.47994
发表时间: 2019-10-01
期刊: ELIFE
影响因子: 7.7
作者:
Graving, Jacob M.;Chae, Daniel;Couzin, Iain D.
通讯作者: Couzin, Iain D.