ITR-(NHS+ASE)-(int+dmc+sim) Automatic Speech Attribute Transcription (ASAT): A Collaborative Speech Research Paradigm and Cyberinfrastructure with Applications to Automatic Speech
ITR-(NHS+ASE)-(int+dmc+sim) Automatic Speech Attribute Transcription (ASAT): A Collaborative Speech Research Paradigm and Cyberinfrastructure with Applications to Automatic Speech
批准号:
0427413
负责人:
Chin-Hui Lee
金额:
$0.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2004
资助国家:
美国
项目状态:
已结题
起止时间:
2004-09-15 至 2011-02-28
中文摘要
长期以来,人们一直假设,人类根据存在于语音知识层次的各个层面(从声学到语用)的检测证据来确定声音的语言身份。事实上,人们不会像自动语音识别(ASR)系统那样不断地将语音信号转换成单词。相反,他们检测声音和听觉证据,权衡它们,并将它们组合起来形成认知假设,然后验证假设,直到达成一致的决定。上述基于人类的语音处理模型为开发下一代语音技术提供了一个候选框架,该技术有可能超越当前的限制。为了弥合ASR系统与人类之间的性能差距,必须将ASR中“语音到文本”的狭隘概念扩展到包含“隐藏”在语音话语中的所有相关人类信息。与传统的自顶向下的网络解码模式不同,我们正在为语音研究建立一个自底向上的事件检测和证据组合范式,以促进协作式自动语音属性转录(ASAT)。本项目的目标是:(1)开发特征检测和知识集成模块,以演示ASAT和ASR;(2)构建开源、高度共享、即插即用的协同研究反卫星网络基础设施,降低反卫星研究的进入门槛;(3)提供客观的评估方法,以监测各个模块和整个系统的技术进步。
英文摘要
It has long been postulated that a human determines the linguistic identity of a sound based on detected evidences that exist at various levels of the speech knowledge hierarchy, from acoustics to pragmatics. Indeed, people do not continuously convert a speech signal into words as an automatic speech recognition (ASR) system attempts to do. Instead, they detect acoustic and auditory evidences, weigh them and combine them to form cognitive hypotheses, and then validate the hypotheses until consistent decisions are reached. The above human-based model of speech processing suggests a candidate framework for developing next generation speech technologies that have the potential to go beyond the current limitations.In order to bridge the performance gap between ASR systems and humans, the narrow notion of speech-to-text in ASR has to be expanded to incorporate all related human information "hidden" in speech utterances. Instead of the conventional top-down, network decoding paradigm for ASR, we are establishing a bottom-up, event detection and evidence combination paradigm for speech research to facilitate collaborative Automatic Speech Attribute Transcription (ASAT). The goals of the proposed project are: (1) develop feature detection and knowledge integration modules to demonstrate ASAT and ASR; (2) build an open source, highly shared, plug-'n'-play ASAT cyberinfrastructure for collaborative research to lower entry barriers to ASR; and (3) provide an objective evaluation methodology to monitor technology advances in individual modules and across the entire system.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SGER: Exploring Universal Acoustic Characterization of Spoken Languages
-
批准号:0639204
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2006
-
负责人:Chin-Hui Lee
-
依托单位:
2003 Symposium on Next Generation Automatic Speech Recognition (ASR)
-
批准号:0352730
-
项目类别:Standard Grant
-
资助金额:$4.96万
-
财政年份:2003
-
负责人:Chin-Hui Lee
-
依托单位:
SGER: Exploring New Auditory Perception Based Approaches to ASR
-
批准号:0350408
-
项目类别:Standard Grant
-
资助金额:$9.95万
-
财政年份:2003
-
负责人:Chin-Hui Lee
-
依托单位:
海外基金