Objective evaluation of the quality of substitution voices

Objective evaluation of the quality of substitution voices
复制标题

客观评价替代声音的质量

DOI:
--
复制
发表时间:
2004
期刊:
European Archives of Oto-Rhino-Laryngology and Head & Neck
影响因子:
--
通讯作者:
P. Dejonckere
P. Dejonckere
中科院分区:
--
文献类型:
--
作者:
M. Moerman;Glenn Pieters;J. Martens;M. V. D. Borgt;P. Dejonckere

文献摘要

参考文献

被引文献

相似文献

本文描述了我们首次尝试开发一种客观评估代声质量的方法。客观分析了表征短声和语音样本的声学参数,如孤立元音序列、VCV和CVCV音节序列、短句等,收集了68例患者(53例全喉切除伴气管食道发音,14例全喉切除伴食道发音,5例行额外侧喉部分切除术)和6例健康人的113例声学参数。每个配准由七个语音话语组成,并接受声学分析和知觉评估,后者涉及“整体印象”、“音调”等八个参数。由于我们工作的目标是找到支持知觉的最佳声学测量并使其精确,因此努力进行基于知觉的声学分析似乎是合乎逻辑的。因此,我们通过带有内置基频(音调)提取器的外围听觉模型进行分析。根据分析器的帧级输出(一帧为10毫秒),针对不同的语音话语计算全局客观参数,例如(1)浊帧的百分比,(2)平均浊音证据,(3)浊音长度分布和(4)基频抖动。为了减少由语音话语的性质(例如,信号中存在停顿、由基音抽取器引起的误差等)引起的参数可变性,使用涉及能量加权和帧选择的非标准平均方案来计算客观参数。对客观参数的统计分析表明,气管食道发音质量优于食道发音,但低于正常发音和保留单侧声带的发音。客观参数和知觉参数之间的相关性是适度的。
This paper describes our first attempts to develop a method for the objective assessment of quality in substitution voices. The objective analysis deals with acoustic parameters characterising short voice and speech samples like a sequence of isolated vowels, a sequence of VCV and CVCVCV syllables, a short sentence, etc. A database of 113 registrations from 68 patients (53 total laryngectomy patients with tracheo-esophageal speech, 14 total laryngectomy patients with esophageal speech and 5 patients with partial frontolateral laryngectomy) and 6 registrations from healthy control persons was collected. Each registration consisted of seven speech utterances and was subjected to an acoustic analysis as well as to a perceptual evaluation, the latter involving eight parameters like “overall impression”, “tonicity”, etc. Since the goal of our work is to find out the best acoustical measurement for supporting perception and making it precise, it seemed logical to strive for a perceptually based acoustic analysis. We therefore performed the analysis by means of a peripheral auditory model with a built-in fundamental frequency (pitch) extractor. From the frame-level outputs (a frame is 10 ms) of the analyser, global objective parameters, such as (1) the percentage of voiced frames, (2) the average voicing evidence, (3) the voicing length distribution and (4) the fundamental frequency jitter, were computed for the different speech utterances. So as to reduce the parameter variability arising from the nature of the speech utterances (e.g., the presence of pauses in the signal, errors caused by the pitch extractor, etc.), the objective parameters were computed using non-standard averaging schemes involving energy weighting and frame selection. A statistical analysis of the objective parameters confirms that the quality of tracheo-esophageal speech is superior to that of esophageal speech, but inferior to that of normal speech and speech with the preservation of one vocal fold. Correlations between the objective parameters and the perceptual parameters are moderate.
DOI: 10.1044/jshr.3601.21
发表时间: 1993-02-01
期刊: JOURNAL OF SPEECH AND HEARING RESEARCH
影响因子: --
作者:
KREIMAN, J;GERRATT, BR;BERKE, GS
通讯作者: BERKE, GS