课题基金 / 基金详情

STIMULATE: Modeling Structure in Speech above the Segment for Spontaneous Speech Recognition

STIMULATE: Modeling Structure in Speech above the Segment for Spontaneous Speech Recognition
刺激:对自发语音识别片段上方的语音结构进行建模
批准号:
9618926
负责人:
Mari Ostendorf
金额:
$45.81万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-03-01 至 1999-09-29

项目摘要

项目成果

Mari Ostendorf的其他基金

相似基金

相关文献

中文摘要
翻译
目前的语音识别技术,虽然在有合作说话者的约束领域很有用,但在无约束的会话或广播语音上仍然导致不可接受的高错误率(30-50%)。这些任务与高精度条件之间的一个重要区别是,即使在单个说话者的数据中,说话风格也有较大的可变性。现有的声学模型没有考虑到这种可变性背后的系统因素,因此必须“更广泛”,导致单词之间更容易混淆,从而导致高错误率。这项工作提出通过在三个时间尺度上表示变异性的来源来改进声学模型:音节、话语中的短区域和说话者。在音节级别,自动聚类将捕捉音节位置和语音缩减效果。在区域层面上,缓慢变化的隐藏说话模式将表明发音的系统性差异,这些差异与发音减少和清晰的语音有关。在说话者层面,语音相关性的分层模型将提高声学模型对少量数据的适应性。实验将涉及会话语音的大词汇识别,使用多通道搜索策略来处理本文提出的高阶模型的成本。通过表示系统的可变性,所提出的工作将显著地推进无约束语音识别的目标任务和更普遍的人机语音通信。
英文摘要
Current speech recognition technology, while useful in constrained domains with cooperative speakers, still leads to unacceptably high error rates (30-50%) on unconstrained conversational or broadcast speech. An important difference between these tasks and high accuracy conditions is the larger variability in speaking style, even within data from a single speaker. Existing acoustic models do not account for the systematic factors behind this variability so must be ``broader,'' leading to more confusability among words and hence high error rates. This work proposes to improve acoustic models by representing sources of variability at three time scales: the syllable, short regions within an utterance, and the speaker. At the syllable level, automatic clustering will capture syllable position and phonetic reduction effects. At the region level, a slowly varying hidden speaking mode will indicate systematic differences in pronunciations associated with reduced vs. clearly articulated speech. At the speaker level, hierarchical models of the correlation among speech sounds will improve adaptation of acoustic models from small amounts of data. Experiments will involve large vocabulary recognition of conversational speech using a multi-pass search strategy to handle the cost of the higher-order models proposed here. By representing systematic variability, the proposed work should significantly advance both the target task of unconstrained speech recognition and human- computer speech communication more generally.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Improving Speech Technology for Better Learning Outcomes: The Case of AAE Child Speakers
  • 批准号:
    2202049
  • 项目类别:
    Standard Grant
  • 资助金额:
    $26.01万
  • 财政年份:
    2022
  • 负责人:
    Mari Ostendorf
  • 依托单位:
RI: Small: Modeling Idiosyncrasies of Speech for Automatic Spoken Language Processing
  • 批准号:
    1617176
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2016
  • 负责人:
    Mari Ostendorf
  • 依托单位:
RI: Small: Simplifying Text for Individual Reading Needs
  • 批准号:
    0916951
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2009
  • 负责人:
    Mari Ostendorf
  • 依托单位:
U.S.-Germany Dissertation Enhancement: Predicting Hidden Structure and Punctuation in Speech for Machine Translation
  • 批准号:
    0552492
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2006
  • 负责人:
    Mari Ostendorf
  • 依托单位:
国内基金
海外基金
Galaxy Analytical Modeling Evolution (GAME) and cosmological hydrodynamic simulations.
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2025
  • 负责人:
    Antonios Katsianis
  • 依托单位: