Relative contributions of acoustic temporal fine structure and envelope cues for lexical tone perception in noise.

Relative contributions of acoustic temporal fine structure and envelope cues for lexical tone perception in noise.
复制标题

DOI:
10.1121/1.4982247
复制
发表时间:
2017-05
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
通讯作者:
B. Qi;Y. Mao;Jiaxing Liu;Bo Liu;Li Xu
B. Qi;Y. Mao;Jiaxing Liu;Bo Liu;Li Xu
中科院分区:
其他
文献类型:
--
作者:
B. Qi;Y. Mao;Jiaxing Liu;Bo Liu;Li Xu

文献摘要

相似文献

以往的研究表明,安静环境下的词汇音调感知依赖于听觉时间精细结构(TFS),而不是信封(E)线索。TFS对噪声环境下语音识别的贡献一直存在争议。在本研究中,普通话声调标记以5种信噪比(-18 ~ +6 dB)混合语音形状噪声(SSN)或双语人牙牙学语(TTB)。然后利用希尔伯特变换分别提取30个波段的TFS和E。从不同信噪比的相同音调标记的声音混合物中创建了25种TFS和E组合。20名听力正常、以普通话为母语的听众参加了声调识别测试。结果表明,无论是TFS还是E的信噪比增加,语音识别性能都有所提高。TTB对音调感知的掩蔽效应弱于SSN。对于两种类型的掩蔽器,TFS和E在噪声中音调感知的感知权重几乎相等,E的作用略大于TFS。因此,TFS和E线索对噪声环境下和假语竞争环境下词汇语调感知的相对贡献不同于安静环境和非声调语言的语音感知。
Previous studies have shown that lexical tone perception in quiet relies on the acoustic temporal fine structure (TFS) but not on the envelope (E) cues. The contributions of TFS to speech recognition in noise are under debate. In the present study, Mandarin tone tokens were mixed with speech-shaped noise (SSN) or two-talker babble (TTB) at five signal-to-noise ratios (SNRs; -18 to +6 dB). The TFS and E were then extracted from each of the 30 bands using Hilbert transform. Twenty-five combinations of TFS and E from the sound mixtures of the same tone tokens at various SNRs were created. Twenty normal-hearing, native-Mandarin-speaking listeners participated in the tone-recognition test. Results showed that tone-recognition performance improved as the SNRs in either TFS or E increased. The masking effects on tone perception for the TTB were weaker than those for the SSN. For both types of masker, the perceptual weights of TFS and E in tone perception in noise was nearly equivalent, with E playing a slightly greater role than TFS. Thus, the relative contributions of TFS and E cues to lexical tone perception in noise or in competing-talker maskers differ from those in quiet and those to speech perception of non-tonal languages.