课题基金 / 基金详情

Deriving Hearing Knowledge from Speech Data

Deriving Hearing Knowledge from Speech Data
从语音数据中获取听力知识
批准号:
1743616
负责人:
Hynek Hermansky
金额:
$7.99万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-06-15 至 2019-05-31

项目摘要

项目成果

Hynek Hermansky的其他基金

相似基金

相关文献

中文摘要
翻译
EArly的探索性研究资助调查了这样一个假设,即语言是为了利用人类的听觉而进化的,因此,人类听觉的特性是印在语言上的。通过优化大量语音数据的语音处理来寻求对这一假设的支持,以便在语音声音之间进行区分。该项目旨在表明,与假设一致的相关听力特性将出现在优化的工程模块中。重点是对更高(皮层)水平的听觉处理进行建模,这在工程项目中通常不被研究。新产生的知识应适用于机器识别的噪声语音。言语中携带的语言信息在时间和频率上都是冗余编码的。冗余,这是在频率上引入的同步道运动和时间上的道惯性,利用人类的认知提取可靠的信息携带的元素从嘈杂的语音。特别地,采用人类听觉的两个特定特性:1)将语音信号的元素分离到不同频率信道中的能力,以及2)提取关于这些信道中的信号的时间动态的信息的能力。特别是,深度神经网络将获取类神经元光谱分析的输出,并在数据上进行训练,以通过一组可学习的二维类皮质光谱-时间滤波器来处理这种类神经元光谱。目前哺乳动物听觉皮层的文献支持这种结构的存在。因此,进展将通过评估与哺乳动物听觉皮层感受野的已知属性的派生2-D过滤器的相似性和它们在提取构成语音消息的潜在语音声音的信息中的有效性来衡量。
英文摘要
This EArly Grant for Exploratory Research investigates the hypothesis that speech evolved to exploit human hearing and, therefore, properties of human hearing are imprinted on speech. A support for this hypothesis is sought by optimizing the speech processing on large amounts of speech data for discrimination among speech sounds. The project intends to show that relevant hearing properties, which are consistent with the hypothesis, will emerge in optimized engineering modules. The focus is on modeling higher(cortical) levels of auditory processing, not usually studied in engineering programs. The new created knowledge should be applicable in machine recognition of noisy speech. Linguistic messages carried in speech are coded redundantly in time and in frequency. Redundancies, which are introduced in frequency by synchronous tract movements and in time by the tract inertia, are exploited by human cognition in extracting reliable information-carrying elements from noisy speech. In particular, two particular properties of human hearing are employed: 1) the ability to separate elements of speech signal into different frequency channels, and 2) the ability to extract information about temporal dynamics of signals in these channels. In particular, a deep neural net would take an output of auditory-like spectral analysis and would be trained on the data to process this auditory-like spectrum through a bank of learnable two-dimensional cortical-like spectro-temporal filters. Existence of such architecture is supported by current literature on mammalian auditory cortex. Therefore, the progress would be gauged by evaluating similarity of the derived 2-D filters with known properties of mammalian auditory cortical receptive fields and by their effectiveness in extracting information about underlying speech sounds that constitute speech messages.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SGER: Multi-Stream Approach Using Syllable Length Temporal Evidence in Acoustic Modeling of Conversational Speech
国内基金
海外基金
基于WHO-HEARING理论框架的老年人听力障碍社区康复模式构建与优化策略研究
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    江帆
  • 依托单位: