The effect of frequency estimation on speech recognition using an acoustic model of a cochlear implant.

The effect of frequency estimation on speech recognition using an acoustic model of a cochlear implant.
复制标题

使用人工耳蜗声学模型进行频率估计对语音识别的影响。

DOI:
10.1016/j.heares.2007.03.002
复制
发表时间:
2007
期刊:
影响因子:
2.8
通讯作者:
Collins,LeslieM
Collins,LeslieM
中科院分区:
医学1区
文献类型:
--
作者:
Throckmorton,ChandraS;SelinKucukoglu,M;Remus,JeremiahJ;Collins,LeslieM

文献摘要

相似文献

在之前的出版物中,研究了通过增加人工耳蜗语音处理算法的每个通道内的频率分辨率来改进语音识别的潜力,作为增加通道数量的替代方案(Throckmorton 等,2006)。这项研究是在正常听力受试者聆听通过语音处理算法的声学模型处理的语音标记的情况下进行的,目的是调查所提出的方法是否改进了语音识别。结果表明,每个通道只需两个离散频率,就可以提高语音识别性能。虽然 Throckmorton 等人(2006)必须考虑并详细讨论了几个注意事项,但这些结果表明,当增加通道数量不是一种选择时,人工耳蜗植入者可能会从通道内增加的频率分辨率中获益。在声学模型中,不是像传统做法那样为每个通道呈现恒定频率(例如 Dorman 等人,1997),而是估计每个通道的频率内容,并根据此估计选择呈现频率。为了计算简单,使用短时傅立叶变换 (STFT) 来分析每个 2 ms 的语音窗口,并在每个通道中选择 N 个预定义呈现频率中最接近的 1 个。然而,考虑到某些通道中的频率内容相对较低,2ms 的窗口持续时间无法对足够的信号进行采样,无法提供准确的频谱估计。作者最初推测,由于无法将刺激速率与精确频率相匹配,这种准确性的缺乏对于人工耳蜗植入来说并不是一个问题,但准确性的缺乏可能会影响正常听力研究的结果(Throckmorton 等,2006)。
In a previous publication, the potential for improving speech recognition by increasing frequency resolution within each channel of a cochlear implant speech processing algorithm was investigated as an alternative to increasing the number of channels (Throckmorton et al., 2006). The study was conducted with normal-hearing subjects listening to speech tokens processed through an acoustic model of a speech processing algorithm with the intent to investigate whether the proposed method improved speech recognition. The results suggested that with as few as two discrete frequencies per channel, an increase in speech recognition performance was possible. Although several caveats must be considered and are discussed at length by Throckmorton et al.(2006), these results suggest the possibility that cochlear implant recipients might derive benefit from increased frequency resolution within channels when increasing the number of channels is not an option.In the acoustic model, rather than presenting a constant frequency for each channel as is traditionally done (eg Dorman et al., 1997), the frequency content of each channel was estimated and a presentation frequency was chosen based on this estimate. For computational simplicity, a short-time Fourier transform (STFT) was used to analyze each 2 ms window of speech and select the closest 1 of N predefined presentation frequencies in each channel. However, given the relatively low frequency content in some of the channels, a 2-ms window duration does not sample enough of the signal to provide an accurate estimate of the frequency spectrum. The authors originally speculated that this lack of accuracy would be less of an issue for implementation in cochlear implants due to the inability to match stimulation rate to an exact frequency, but that the lack of accuracy may have influenced results in the normal-hearing study (Throckmorton et al., 2006).