课题基金 / 基金详情

CAREER: Landmark-Based Speech Recognition in Music and Speech Backgrounds

CAREER: Landmark-Based Speech Recognition in Music and Speech Backgrounds
职业:音乐和语音背景中基于地标的语音识别
批准号:
0132900
负责人:
Mark Hasegawa-Johnson
金额:
$39.58万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2002
资助国家:
美国
项目状态:
已结题
起止时间:
2002-07-01 至 2008-06-30

项目摘要

项目成果

Mark Hasegawa-Johnson的其他基金

相似基金

相关文献

中文摘要
翻译
这是一个教师早期职业发展(Career)奖。该研究将开发语音识别和听觉场景分析模型,这些模型是概率分布,其参数可以从数据中训练出来,其内部结构能够抽象出人类听众的感知反应模式。将探讨两个广泛的研究问题:(1)代表声源的音高、包络和时间的概率模型能否以一种易于处理的方式计算和集成?(2)基于地标声学特征的概率模型划分、训练和识别评分的理论和经验要求是什么?语音中的标志是随着时间的推移,语音流中可识别的点,如辅音的释放和闭合,元音的中心和滑音的极端。这个项目的教育部分包括在本科和研究生阶段的重要课程开发,以及对本科生和研究生研究受训者的指导的大力投资。该奖项旨在表彰和支持有可能成为21世纪学术领袖的教师学者的早期职业发展活动。这是声学和计算机科学的基础科学研究,但它解决了一个非常实际的问题,即计算机在识别语音方面仍然比人类差得多。语音识别技术已经成为一个重要的产业,但它在未来会变得更加重要,因为移动计算和以计算机为媒介的通信使得数以百万计的人有必要口头而不是通过键盘来控制机器。这项工作的教育部分将培养研究生成为教师和传播者以及研究人员,从而使他们为帮助建立这个令人兴奋的、不断发展的领域所需的人才基础做好准备。
英文摘要
This is a Faculty Early Career Development (CAREER) award. The research will develop speech recognition and auditory scene analysis models that are probability distributions whose parameters can be trained from data and whose internal structures are capable of abstracting the perceptual response patterns of human listeners. Two broad research questions will be explored: (1) Can probability models representing the pitch, envelope, and timing of an acoustic source be computed and integrated in a tractable manner? (2) What are the theoretical and empirical requirements for the partitioning, training, and recognition scoring of probability models for landmark-based acoustic features? Landmarks in speech are identifiable points in the flow of sound over time, such as consonant releases and closures, vowel centers, and glide extrema. The educational component of this project includes significant curriculum development at both the undergraduate and graduate levels, and a strong investment in the mentoring of undergraduate and graduate research trainees.This CAREER award recognizes and supports the early career-development activities of a teacher-scholar who is likely to become an academic leader of the twenty-first century. This is fundamental scientific research in acoustics and computer science, but it addresses the very practical problem that computers are still far worse at recognizing speech than human beings are. Speech recognition technology has already become an important industry, but it will become far more important in the future as mobile computing and computer-mediated communications make it necessary for millions of people to control machines verbally rather than by means of keyboards. The educational component of this work will train graduate students to be teachers and communicators, as well as researchers, thus preparing them to help build the base of personnel needed in this exciting, growing area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
FAI: A New Paradigm for the Evaluation and Training of Inclusive Automatic Speech Recognition
RI: Small: Collaborative Research: Automatic Creation of New Speech Sound Inventories
EAGER: Matching Non-Native Transcribers to the Distinctive Features of the Language Transcribed
FODAVA-Partner: Visualizing Audio for Anomaly Detection
国内基金
海外基金
基于Landmark知识的规划方法研究
  • 批准号:
    61103136
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    22.0万元
  • 批准年份:
    2011
  • 负责人:
    蔡敦波
  • 依托单位: