STIMULATE: Modeling and Automatic Labeling of Hidden Word- Level Events in Spontaneous Speech
STIMULATE: Modeling and Automatic Labeling of Hidden Word- Level Events in Spontaneous Speech
批准号:
9619921
负责人:
Elizabeth Shriberg
金额:
$77.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-03-01 至 2006-02-28
中文摘要
大多数当前的自然语言处理技术期望类似于阅读或受限语音的输入。当应用于自发演讲时,这种技术会遇到两个严重的困难。首先,自发言语包含与输入的非命题方面相关的表面现象,如不流利和话语标记语。其次,自发讲话缺乏明显的标点符号,无法将输入分割成有意义的单元。对于有效的自然语言处理,这些现象应该在输入中公开标记;然而,目前的语音识别器只产生原始的单词序列。该项目的目标是增强语音识别模型,以输出为这些现象注释的单词序列,称为“隐藏词级事件”(HWES)。开发了新的模型,以允许HWE的识别与单词识别同时进行。HWE识别基于声学和语言模型的组合,扩展了当前系统中的标准组件。新模型还捕捉了HWE的韵律特征,包括语调和时长模式。韵律信息与描述HWE相对于词汇和句法单位的分布的统计语言模型相结合。研究结果将显著提高我们自动处理自发言语的能力;这项研究也将加深我们对自发言语的基本理解。
英文摘要
Most current NLP techniques expect input resembling read or constrained speech. When applied to spontaneous speech, such techniques encounter two serious difficulties. First, spontaneous speech contains surface phenomena relating to non-propositional aspects of the input, such as disfluencies and discourse markers. Second, spontaneous speech lacks overt punctuation for segmenting the input into meaningful units. For effective NLP, such phenomena should be overtly marked in the input; current speech recognizers, however, produce only a raw sequence of words. The goal of this project is to augment speech recognition models to output word sequences annotated for these phenomena, termed "Hidden Word-Level Events" (HWEs). New models are developed to allow recognition of HWEs to occur in tandem with word recognition. HWE recognition is based on a combination of acoustic and language models, extending the standard components found in current systems. The new models also capture prosodic characteristics of HWEs, including intonation and duration patterns. The prosodic information is combined with statistical language models describing the distribution of HWEs in relation to lexical and syntactic units. Results should significantly enhance our ability to process spontaneous speech automatically; the research will also further our basic understanding of spontaneous speech.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: A Corpus of Aligned Speech and ANS Sensor Data
-
批准号:1449202
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2014
-
负责人:Elizabeth Shriberg
-
依托单位:
TalkPrinting: New Features and Models for Automatic Speaker Recognition
-
批准号:0544682
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2005
-
负责人:Elizabeth Shriberg
-
依托单位:
Modeling Disfluencies in Spontaneous Speech
-
批准号:9314967
-
项目类别:Continuing Grant
-
资助金额:$68.97万
-
财政年份:1994
-
负责人:Elizabeth Shriberg
-
依托单位:
NSF-NATO Postdoctoral Fellowhips
-
批准号:9353732
-
项目类别:Fellowship Award
-
资助金额:$0.0万
-
财政年份:1993
-
负责人:Elizabeth Shriberg
-
依托单位:
国内基金
海外基金
Galaxy Analytical Modeling
Evolution (GAME) and cosmological
hydrodynamic simulations.
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2025
-
负责人:Antonios Katsianis
-
依托单位: