Importance of tonal envelope cues in Chinese speech recognition.

Importance of tonal envelope cues in Chinese speech recognition.
复制标题

DOI:
10.1121/1.413004
复制
发表时间:
1998
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
通讯作者:
Qian-Jie Fua
Qian-Jie Fua
中科院分区:
其他
文献类型:
--
作者:
Qian-Jie Fua

文献摘要

被引文献

相似文献

近年来的研究表明,时间波形包络线索可以为英语语音识别提供重要的信息。本研究调查了声调语言普通话中时间包络线索的使用。在本研究中,将语音划分为几个频率分析波段;通过半波整流和低通滤波提取各波段的幅值包络,用于调制与分析波段相同带宽的噪声。这些操作保留了每个频带的时间和幅度线索,但删除了每个频带内的频谱细节。12名母语为汉语的听者分别用1、2、3、4个噪声带识别汉语元音、辅音、声调和句子。结果表明,元音、辅音和句子的识别得分随频带数的增加而单调增加,与英语语音识别的模式相似。相比之下,音调的识别准确率始终保持在80%左右,与频带的数量无关。这种高水平的音调识别使得汉语和英语在无谱信息的单波段条件下的开放集句子识别差异显著(11.0%)。数据还显示,在主要的时间线索下,降调(音调3)和降调(音调4)比平调(音调1)和升调(音调2)更容易被识别。音调识别中的这种差异模式导致了单词识别中的类似模式:音调为3或4的单词更容易被识别,而音调为1和2的单词则不容易被识别。利用幂函数模型进一步探讨了声调在汉语语音识别中的定量作用,发现声调在音素识别和句子识别之间起着重要的作用。
Recent studies have shown that temporal waveform envelope cues can provide significant information for English speech recognition. This study investigated the use of temporal envelope cues in a tonal language: Mandarin Chinese. In this study, the speech was divided into several frequency analysis bands; the amplitude envelope was extracted from each band by half-wave rectification and low-pass filtering and was used to modulate a noise of the same bandwidth as the analysis band. These manipulations preserved temporal and amplitude cues in each frequency band, but removed the spectral detail within each band. Chinese vowels, consonants, tones and sentences were identified by 12 native Chinese-speaking listeners with 1, 2, 3, and 4 noise bands. The results showed that the recognition score of vowels, consonants, and sentences increased monotonically with the number of bands, a pattern similar to that observed in English speech recognition. In contrast, tones were consistently recognized at about 80% correct level, independent of the number of bands. This high level of tone recognition produced a significant difference in the open-set sentence recognition between Chinese (11.0%) and English (2.9%) for the one-band condition where no spectral information was available. The data also revealed that, with primarily temporal cues, the falling-rising tone (tone 3) and the falling tone (tone 4) were more easily recognized than the flat tone (tone 1) and the rising tone (tone 2). This differential pattern in tone recognition resulted in a similar pattern in word recognition: words having either tone 3 or 4 were more likely to be recognized while words having tone 1 and 2 were not. The quantitative role of tones in Chinese speech recognition was further explored using a power-function model and found to play a significant role in relating phoneme recognition to sentence recognition.