课题基金 / 基金详情

Modeling Pronunciation Variation for Universal Access to Speech Understanding

Modeling Pronunciation Variation for Universal Access to Speech Understanding
为普遍获得语音理解而建模发音变化
批准号:
9978025
负责人:
Daniel Jurafsky
金额:
$50.4万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1999
资助国家:
美国
项目状态:
已结题
起止时间:
1999-09-15 至 2004-08-31

项目摘要

项目成果

Daniel Jurafsky的其他基金

相似基金

相关文献

中文摘要
翻译
为了让基于语音的界面普遍可用--即所有公民在所有情况下都可以使用--它们必须成功地处理发音的巨大差异,因为研究表明,这是迄今为止机器语音识别中最大的错误来源。这个项目将通过建立动态发音模型来解决这个问题,该模型使用上下文来预测一个单词在特定话语中最可能的发音。PIs以前的研究已经确定,在各种情况下,说话者更有可能删除某些声音(例如,“About”中的最后一个“t”):如果下一个单词以某些辅音开头;如果他们说得很快;如果这个单词并不令人惊讶(具有很高的三元组概率);如果他们是年轻人和男性;如果这个单词后面没有停顿或重复单词;或者如果他们说某些方言。他们进一步展示了如何使用置信度度量来根据语音中的实际发音来识别词典中不准确的发音。这个项目将扩展和结合这两个研究领域,通过使用语音置信度模型来识别导致单词错误的发音错误的单词,然后为这些单词建立发音变化的统计(决策树)模型。PI的模型提供了一个通用引擎,它们将应用于嵌入Sphinx-II语音识别器中的几种特定类别的发音变化,包括方言和说话人类型。计算模型在发音变异研究中的应用是语言学和认知建模领域的一个重要的新研究领域,对科学和教育学的许多领域都有影响和潜在的好处,包括普遍获得。
英文摘要
In order for speech-based interfaces to be universally accessible - that is, available to all citizens in all situations - they must deal successfully with massive variability in pronunciation, as studies have shown that this is by far the single largest source of error in machine speech recognition. This project will address that issue by building dynamic pronunciation models which use context to predict the most likely pronunciations of a word in a particular utterance. The PIs' previous work has established that speakers are more likely to delete certain sounds (e.g., the final "t" in "about") in a variety of situations: if the next word starts with certain consonants; if they are speaking quickly; if the word is unsurprising (has a high trigram probability); if they are young and male; if the word is not followed by a pause or a word-repetition; or if they are speakers of certain dialects. They have further shown how confidence metrics can be used to identify inaccurate dictionary pronunciations based on actual pronunciations in speech. This project will extend and combine both of these lines of research, by using phonetic-confidence models to identify words with incorrect pronunciations that contribute to word-error, and then building statistical (decision-tree) models of pronunciation variation for these words. The PIs' model provides a general engine which they will apply to several specific classes of pronunciation variation, including dialect and speaker type, embedded in the Sphinx-II speech recognizer. The application of computational models to the study of pronunciation variation is an important new research area in linguistics and cognitive modeling, with implications for and potential benefits to many areas of science and pedagogy, including universal access.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: New tools for studying structural and inductive bias in NLP models
  • 批准号:
    2128145
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2021
  • 负责人:
    Daniel Jurafsky
  • 依托单位:
RI: Medium: Deep Understanding: Integrating Neural and Symbolic Models of Meaning
  • 批准号:
    1514268
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $110.0万
  • 财政年份:
    2015
  • 负责人:
    Daniel Jurafsky
  • 依托单位:
RI: Small: Learning Meaning and Grammar from Interaction, Context, and the World
  • 批准号:
    1216875
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2012
  • 负责人:
    Daniel Jurafsky
  • 依托单位:
RI-Small: Unsupervised Learning of Meaning
  • 批准号:
    0811974
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2008
  • 负责人:
    Daniel Jurafsky
  • 依托单位:
海外基金