Predicting Speech Perception in Older Listeners with Sensorineural Hearing Loss Using Automatic Speech Recognition

Predicting Speech Perception in Older Listeners with Sensorineural Hearing Loss Using Automatic Speech Recognition
复制标题

DOI:
10.1177/2331216520914769
复制
发表时间:
2020-03-01
期刊:
影响因子:
2.7
通讯作者:
Fullgrabe, Christian
Fullgrabe, Christian
中科院分区:
医学1区
文献类型:
--
作者:
Fontan, Lionel;Cretin-Maitenaz, Tom;Fullgrabe, Christian

文献摘要

被引文献

相似文献

本研究的目的是提供一个概念证明,在安静的无辅助老年听力受损(OHI)的听众的语音清晰度可以预测的自动语音识别(ASR)。24名OHI听众使用不同语言复杂性和可预测性的语音材料完成了三项语音识别任务(即,logatoms、单词和句子)。ASR系统首先在不同的语音材料上进行训练,然后用于识别呈现给听众的相同语音刺激,但经过处理以模仿每个听众所经历的与年龄相关的听力损失的一些感知后果:听力阈值的提高(通过线性滤波)、频率选择性的损失(通过频谱拖尾)和响度恢复(通过将幅度包络提高到幂)。独立于ASR系统中使用的词汇的大小,在人类和机器可懂度分数之间观察到强到非常强的相关性。然而,大的均方根误差(RMSE)观察到所有条件。频率选择性损失的模拟对相关性的强度和RMSE有负面影响。最高的相关性和最小的RMSE被发现为logatoms,这表明预测系统主要反映了听觉系统的外围部分的功能。在句子的情况下,考虑到认知表现,人类可懂度的预测显着提高。这项研究首次表明,ASR,即使在完整的独立语音材料的训练,可以用来估计OHI听众的语音清晰度的趋势。
The objective of this study was to provide proof of concept that the speech intelligibility in quiet of unaided older hearing-impaired (OHI) listeners can be predicted by automatic speech recognition (ASR). Twenty-four OHI listeners completed three speech-identification tasks using speech materials of varying linguistic complexity and predictability (i.e., logatoms, words, and sentences). An ASR system was first trained on different speech materials and then used to recognize the same speech stimuli presented to the listeners but processed to mimic some of the perceptual consequences of age-related hearing loss experienced by each of the listeners: the elevation of hearing thresholds (by linear filtering), the loss of frequency selectivity (by spectrally smearing), and loudness recruitment (by raising the amplitude envelope to a power). Independently of the size of the lexicon used in the ASR system, strong to very strong correlations were observed between human and machine intelligibility scores. However, large root-mean-square errors (RMSEs) were observed for all conditions. The simulation of frequency selectivity loss had a negative impact on the strength of the correlation and the RMSE. Highest correlations and smallest RMSEs were found for logatoms, suggesting that the prediction system reflects mostly the functioning of the peripheral part of the auditory system. In the case of sentences, the prediction of human intelligibility was significantly improved by taking into account cognitive performance. This study demonstrates for the first time that ASR, even when trained on intact independent speech material, can be used to estimate trends in speech intelligibility of OHI listeners.