课题基金 / 基金详情

Development of innovative speech enhancement algorithms based on the central auditory system.

Development of innovative speech enhancement algorithms based on the central auditory system.
开发基于中央听觉系统的创新语音增强算法。
批准号:
RGPIN-2014-05301
负责人:
Plourde, Eric
金额:
$1.6万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2016
资助国家:
加拿大
项目状态:
已结题
起止时间:
2016-01-01 至 2017-12-31

项目摘要

项目成果

Plourde, Eric的其他基金

相似基金

相关文献

中文摘要
翻译
数百万加拿大人通常在嘈杂的环境中使用平板电脑、智能手机以及现在的智能手表和眼镜等多媒体设备。这些设备包括许多语音处理算法,如语音编码器或自动语音识别器(ASR),其性能受到噪声的严重影响。例如,ASR可以在无噪声环境中正确识别85%的单词;但是,在信噪比(SNR)为10分贝的情况下,这一百分比可能会下降到31%。为了限制这种性能下降,在这些设备中包括语音增强(SE)模块,其目的是在不影响语音质量的情况下降低噪声水平。这些SE模块的性能在很大程度上不是最优的。事实上,一项比较了14种最好的SE算法的研究报告称,在信噪比为10分贝的情况下,最大主观得分仅为3/5。与之形成鲜明对比的是,听觉系统能很好地处理噪音。事实上,在相对嘈杂的环境中,人类很容易跟上对话。 因此,我的研究计划的长期目标(+10年)是开发商业上可行的SE算法,其灵感来自于中央听觉系统,即位于听神经和听觉皮质之间的听觉系统的一部分,目标是接近听觉系统在存在噪声的情况下的出色性能。在短期内(不到5年),主要目标将是对中枢听觉系统神经元对噪声发声的表征进行统计建模,并在这些模型的基础上开发SE算法。为了实现这些短期目标,我们将首先将神经放电表示为点过程,并使用这种表示来建立神经编码和解码的统计模型;神经编码是在给定刺激的情况下对尖峰序列的估计,例如发声,神经解码是对给定尖峰序列的刺激的估计。这些模型将特别考虑发声过程中噪声的存在。此外,我们将使用导出的模型来开发SE的统计估计器。由于这些统计估计器将被设置在更接近中枢听觉系统的一个域中,我们期望得到的估计器在感知上更相关,从而更有效。 神经信号的精确而简单的统计模型的最新发展为SE开辟了一条很有前途的研究途径,将在当前的提案中加以利用。此外,这项提议将使培训具有神经科学、统计信号处理和语音处理技能的多学科研究人员成为可能。该计划完成后,将提高SE模块的性能,因此将允许更有效地使用数百万台多媒体便携式设备,如平板电脑、智能手机、手表或眼镜。
英文摘要
Multimedia devices such as tablets, smart phones and now smart watches and glasses are commonly used in noisy environments by millions of Canadians. These devices include many speech processing algorithms, such as speech coders or automatic speech recognizers (ASR), whose performances are seriously affected by the presence of noise. For example, an ASR can identify 85% of the words correctly in a noise-free environment; however, this percentage can drop to 31% with a signal-to-noise ratio (SNR) of 10 dB. In order to limit this decrease in performance, speech enhancement (SE) modules, which aim at reducing the noise level without affecting the speech quality, are included in these devices. The performance of these SE modules is largely sub-optimal. In fact, a study having compared 14 of the best SE algorithms reports a maximum subjective score of barely 3/5 for an SNR of 10 dB. In sharp contrast, the auditory system deals very well with noise. In fact, it is fairly easy for humans to follow a conversation in a relatively noisy environment. The long-term objective (+ 10 years) of my research program is thus to develop commercially viable SE algorithms that are inspired by the central auditory system, i.e. the part of the auditory system between the auditory nerve and auditory cortex, with the goal of approaching the excellent performance of the auditory system in the presence of noise. In the short-term (less than 5 years), the main objectives will be to statistically model the representation of noisy vocalizations by the neurons of the central auditory system as well as to develop SE algorithms based on these models. To achieve these short-term objectives, we will first represent neural discharges as point processes and use this representation to develop statistical models of neural coding and decoding; neural coding being the estimation of a spike train given a stimulus, such as a vocalization, and neural decoding, the estimation of a stimulus given a spike train. These models will specifically take into account the presence of noise in vocalizations. Furthermore, we will use the derived models to develop statistical estimators for SE. Since these statistical estimators will be set in a domain closer to the one of the central auditory system, we expect the resulting estimators to be more perceptually relevant and thus more efficient. The recent development of accurate, yet simple, statistical models of neural signals opens a promising research avenue for SE that will be exploited in the current proposal. Moreover, this proposal will allow for the training of multidisciplinary researchers having skills in neuroscience, statistical signal processing and speech processing. Upon completion, this program will improve the performance of SE modules and will therefore allow a much more efficient use of millions of multimedia portable devices such as tablets, smart phones, watches or glasses.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Design and implementation of spiking neural network based speech enhancement algorithms
  • 批准号:
    RGPIN-2020-05077
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.4万
  • 财政年份:
    2022
  • 负责人:
    Plourde, Eric
  • 依托单位:
Design and implementation of spiking neural network based speech enhancement algorithms
  • 批准号:
    RGPAS-2020-00112
  • 项目类别:
    Discovery Grants Program - Accelerator Supplements
  • 资助金额:
    $2.91万
  • 财政年份:
    2022
  • 负责人:
    Plourde, Eric
  • 依托单位:
Design and implementation of spiking neural network based speech enhancement algorithms
  • 批准号:
    RGPAS-2020-00112
  • 项目类别:
    Discovery Grants Program - Accelerator Supplements
  • 资助金额:
    $2.91万
  • 财政年份:
    2021
  • 负责人:
    Plourde, Eric
  • 依托单位:
Design and implementation of spiking neural network based speech enhancement algorithms
  • 批准号:
    RGPIN-2020-05077
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.4万
  • 财政年份:
    2021
  • 负责人:
    Plourde, Eric
  • 依托单位:
海外基金