课题基金 / 基金详情

ITR: Dynamics-based Speech Segregation

ITR: Dynamics-based Speech Segregation
ITR:基于动力学的语音分离
批准号:
0081058
负责人:
DeLiang Wang
金额:
$45.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2000
资助国家:
美国
项目状态:
已结题
起止时间:
2000-09-01 至 2004-08-31

项目摘要

项目成果

DeLiang Wang的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
A typical auditory scene contains multiple simultaneous events, and a remarkable feat of the auditory nervous system is its ability to disentangle the acoustic mixture and group the acoustic energy from the same event. This fundamental process of auditory perception is called auditory scene analysis. Of particular importance in auditory scene analysis is the separation of speech from interfering sounds, or speech segregation. Speech segregation remains a largely unsolved problem in auditory engineering and speech technology. In this project, the P1 seeks to develop a dynamics-based system for speech segregation using perceptual and neural principles. Auditory grouping will be based on oscillatory correlation, whereby phases of neural oscillators encode the binding of auditory features. The investigation will consist of subsequent stages of computation, starting from simulated auditory periphery composed of cochlear filtering and hair cell transduction. A mid-level representation will be formed by computing auto- and cross-correlation of filter channels. A stage of segment formation then creates individual elements of a represented auditory scene, each of which is a dynamically evolving, connected time-frequency structure that may overlap with other elements. Operating on auditory segments from the segment formation stage, both simultaneous organization and sequential organization will be incorporated. For simultaneous organization, grouping will be based on periodicity, location, onset and offset analyses, while for sequential organization grouping will be based on pitch, spectral, and location continuities. In particular, two pitch maps corresponding to two ears and one location map will be computed for auditory organization. All of the employed grouping cues are consistent with perceptual principles of auditory scene analysis. These cues guide the connectivity of neural oscillator networks, which perform grouping and segregation of auditory segments. The proposed system will be evaluated using real recordings of speech and interfering sounds, where speech can be both voiced and unvoiced. The success of the system will be quantitatively assessed using two measures: changes in signal-to-noise ratio and speech recognition rate. This project is expected to make significant contributions to automatic speech recognition in unconstrained environments.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Deep neural networks for multi-channel speaker localization and speech separation
  • 批准号:
    1808932
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2018
  • 负责人:
    DeLiang Wang
  • 依托单位:
Collaborative Research: Separating Speech from Speech Noise to Improve Speech Intelligibility
Automated Auditory Scene Analysis Based on Oscillatory Correlation
Segmentation and Recognition of Complex Temporal Patterns
国内基金
海外基金
β-arrestin2- MFN2-Mitochondrial Dynamics轴调控星形胶质细胞功能对抑郁症进程的影响及机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2023
  • 负责人:
  • 依托单位: