ITR-(NHS+ASE)-(int+dmc+sim) Automatic Speech Attribute Transcription (ASAT): A Collaborative Speech Research Paradigm and Cyberinfrastructure with Applications to Automatic Speech
ITR-(NHS+ASE)-(int+dmc+sim) Automatic Speech Attribute Transcription (ASAT): A Collaborative Speech Research Paradigm and Cyberinfrastructure with Applications to Automatic Speech
批准号:
0427413
负责人:
Chin-Hui Lee
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-09-15 至 2011-02-28
中文摘要
长期以来,人们一直假设人类根据存在于从声学到语用学的语音知识层次的各个层面上检测到的证据来确定声音的语言身份。事实上,人们并不像自动语音识别(ASR)系统试图做的那样,连续地将语音信号转换成单词。取而代之的是,他们检测声学和听觉证据,对它们进行权衡,并将它们结合起来形成认知假设,然后验证这些假设,直到得出一致的决定。上述基于人类的语音处理模型为开发下一代语音技术提供了一个候选框架,该框架具有超越当前限制的潜力。为了弥合ASR系统和人类之间的性能差距,ASR中狭隘的语音到文本的概念必须扩展到包含所有隐藏在语音话语中的相关人类信息。我们正在建立一种自下而上、事件检测和证据组合的语音研究范式,以促进协同自动语音属性转录(ASAT),而不是传统的自上而下、网络解码的ASR范式。拟议项目的目标是:(1)开发特征检测和知识集成模块,以展示反卫星技术和反卫星技术;(2)建立一个开放源码、高度共享、即插即用的反卫星技术网络基础设施,用于协作研究,以降低进入反卫星技术的门槛;(3)提供客观的评价方法,以监测各个模块和整个系统的技术进步。
英文摘要
It has long been postulated that a human determines the linguistic identity of a sound based on detected evidences that exist at various levels of the speech knowledge hierarchy, from acoustics to pragmatics. Indeed, people do not continuously convert a speech signal into words as an automatic speech recognition (ASR) system attempts to do. Instead, they detect acoustic and auditory evidences, weigh them and combine them to form cognitive hypotheses, and then validate the hypotheses until consistent decisions are reached. The above human-based model of speech processing suggests a candidate framework for developing next generation speech technologies that have the potential to go beyond the current limitations.In order to bridge the performance gap between ASR systems and humans, the narrow notion of speech-to-text in ASR has to be expanded to incorporate all related human information "hidden" in speech utterances. Instead of the conventional top-down, network decoding paradigm for ASR, we are establishing a bottom-up, event detection and evidence combination paradigm for speech research to facilitate collaborative Automatic Speech Attribute Transcription (ASAT). The goals of the proposed project are: (1) develop feature detection and knowledge integration modules to demonstrate ASAT and ASR; (2) build an open source, highly shared, plug-'n'-play ASAT cyberinfrastructure for collaborative research to lower entry barriers to ASR; and (3) provide an objective evaluation methodology to monitor technology advances in individual modules and across the entire system.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SGER: Exploring Universal Acoustic Characterization of Spoken Languages
-
批准号:0639204
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2006
-
负责人:Chin-Hui Lee
-
依托单位:
2003 Symposium on Next Generation Automatic Speech Recognition (ASR)
-
批准号:0352730
-
项目类别:Standard Grant
-
资助金额:$4.96万
-
财政年份:2003
-
负责人:Chin-Hui Lee
-
依托单位:
SGER: Exploring New Auditory Perception Based Approaches to ASR
-
批准号:0350408
-
项目类别:Standard Grant
-
资助金额:$9.95万
-
财政年份:2003
-
负责人:Chin-Hui Lee
-
依托单位:
海外基金