Speaker recognition with temporal cues in acoustic and electric hearing.

Speaker recognition with temporal cues in acoustic and electric hearing.
复制标题

DOI:
10.1121/1.1944507
复制
发表时间:
2005-08
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
通讯作者:
Michael Vongphoe;F. Zeng
Michael Vongphoe;F. Zeng
中科院分区:
其他
文献类型:
--
作者:
Michael Vongphoe;F. Zeng

文献摘要

被引文献

相似文献

自然口语处理不仅包括语音识别,还包括对说话人的性别、年龄、情感和社会地位的识别。我们在这项研究中的目的是评估时间线索是否足以支持语音和说话人识别。10名植入人工耳蜗的受试者和6名听力正常的受试者被出示了由3名男性、3名女性、2名男孩和2名女孩说出的元音符号。在一种情况下,受试者被要求识别元音。在另一种情况下,受试者被要求指认说话人。为说话人识别任务提供了广泛的训练。听力正常的受试者在这两项任务中都取得了近乎完美的表现。人工耳蜗受试者在元音识别方面表现良好,但在说话人识别方面表现不佳。人工耳蜗的功能水平与正常水平相当,有8个谱带用于元音识别,但只有1个谱带用于说话人识别。这些结果显示语音和说话人识别之间的分离主要与时间线索有关,突出了当前人工耳蜗语音处理策略的局限性。针对当前人工耳蜗使用者的语音识别问题,提出了基频显式编码和调频等方法。
Natural spoken language processing includes not only speech recognition but also identification of the speaker's gender, age, emotional, and social status. Our purpose in this study is to evaluate whether temporal cues are sufficient to support both speech and speaker recognition. Ten cochlear-implant and six normal-hearing subjects were presented with vowel tokens spoken by three men, three women, two boys, and two girls. In one condition, the subject was asked to recognize the vowel. In the other condition, the subject was asked to identify the speaker. Extensive training was provided for the speaker recognition task. Normal-hearing subjects achieved nearly perfect performance in both tasks. Cochlear-implant subjects achieved good performance in vowel recognition but poor performance in speaker recognition. The level of the cochlear implant performance was functionally equivalent to normal performance with eight spectral bands for vowel recognition but only to one band for speaker recognition. These results show a disassociation between speech and speaker recognition with primarily temporal cues, highlighting the limitation of current speech processing strategies in cochlear implants. Several methods, including explicit encoding of fundamental frequency and frequency modulation, are proposed to improve speaker recognition for current cochlear implant users.