Temporal Segmentation for Laryngeal High-Speed Videoendoscopy in Connected Speech.

Temporal Segmentation for Laryngeal High-Speed Videoendoscopy in Connected Speech.
复制标题

DOI:
10.1016/j.jvoice.2017.05.014
复制
发表时间:
2018-03
期刊:
Journal of voice : official journal of the Voice Foundation
影响因子:
--
通讯作者:
Orlikoff RF
Orlikoff RF
中科院分区:
其他
文献类型:
--
作者:
Naghibolhosseini M;Deliyski DD;Zacharias SRC;de Alarcon A;Orlikoff RF

文献摘要

参考文献

被引文献

相似文献

本研究提出了一种基于梯度的方法,用于喉高速视频内窥镜(HSV)数据的时间分割在连接的语音。一个定制开发的HSV系统加上灵活的光纤鼻喉镜被用来记录一个声音正常的女性参与者在阅读的“彩虹通道”。提出了一种基于梯度的运动窗口生成算法。当应用于HSV数据时,运动窗口充当跟踪振动声带位置的滤波器。声门面积的波形估计使用基于声学的图像处理方法。声带振动频率的计算是基于自相关提取的基本频率(f0)从声门区波形。然后基于f0轮廓和会厌阻塞的自动检测进行时间分割。此外,通过逐帧查看HSV图像来执行视觉时间分割,以确定发声起始和偏移的时间点以及声门的会厌阻塞。对自动和视觉时间分割方法产生的时间点进行交叉验证。发现自动算法产生的上升和下降的f0轮廓模式与HSV数据中振动频率变化的目视检查一致。本研究证明了连接语音的HSV成像的自动时间分割的可行性,其允许将视频内容映射到每个发声的起始、偏移和会厌障碍。自动分析HSV成像的连接语音具有显着的临床潜力,推进仪器语音评估协议。
This study proposes a gradient-based method for temporal segmentation of laryngeal high-speed videoendoscopy (HSV) data obtained during connected speech. A custom-developed HSV system coupled with a flexible fiberoptic nasolaryngoscope was used to record one vocally normal female participant during reading of the “Rainbow Passage.” A gradient-based algorithm was developed to generate a motion window. When applied to the HSV data, the motion window acted as a filter tracking the location of the vibrating vocal folds. The glottal area waveform was estimated using a statistical-based image-processing approach. The vocal fold vibratory frequency was computed by an autocorrelation-based extraction of the fundamental frequency (f0) from the glottal area waveform. Temporal segmentation was then performed based on the f0 contour and automatic detection of the epiglottic obstructions. Additionally, visual temporal segmentation was performed by viewing the HSV images frame by frame to determine the time points of the vocalization onsets and offsets, and the epiglottic obstructions of the glottis. The time points resulting from the automatic and visual temporal segmentation methods were cross-validated. The f0-contour patterns of rise and fall resulting from the automatic algorithm were found to be in agreement with the visual inspection of the vibratory frequency change in the HSV data. This study demonstrated the feasibility of automatic temporal segmentation of HSV imaging of connected speech, which allows for mapping the video content into onsets, offsets, and epiglottic obstructions for each vocalization. Automated analysis of HSV imaging of connected speech has significant clinical potential for advancing instrumental voice assessment protocols.
DOI: 10.1159/000111802
发表时间: 2008-01-01
影响因子: 1
作者:
Deliyski, Dimitar D.;Petrushev, Pencho P.;Hillman, Robert E.
通讯作者: Hillman, Robert E.
DOI: 10.1097/01.mlg.0000154739.48314.ee
发表时间: 2005-02-01
期刊: LARYNGOSCOPE
影响因子: 2.6
作者:
Roy, N;Gouse, M;Smith, ME
通讯作者: Smith, ME
DOI: 10.1159/000077798
发表时间: 2004-01-01
期刊: ORL-JOURNAL FOR OTO-RHINO-LARYNGOLOGY AND ITS RELATED SPECIALTIES
影响因子: --
作者:
Halberstam, B
通讯作者: Halberstam, B
DOI: 10.1097/moo.0b013e3282fe96ce
发表时间: 2008-06-01
影响因子: 1.6
作者:
Mehta, Daryush D.;Hillman, Robert E.
通讯作者: Hillman, Robert E.
DOI: 10.1097/moo.0b013e3283395dd4
发表时间: 2010-06
影响因子: 1.6
作者:
Deliyski DD;Hillman RE
通讯作者: Hillman RE