Modeling Pronunciation Variation for Universal Access to Speech Understanding
Modeling Pronunciation Variation for Universal Access to Speech Understanding
批准号:
9978025
负责人:
Daniel Jurafsky
金额:
$50.4万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1999
资助国家:
美国
项目状态:
已结题
起止时间:
1999-09-15 至 2004-08-31
中文摘要
为了让基于语音的界面能够被普遍使用——也就是说,所有公民在任何情况下都可以使用——他们必须成功地处理语音的巨大变化,因为研究表明,这是迄今为止机器语音识别中最大的错误来源。这个项目将通过建立动态发音模型来解决这个问题,该模型使用上下文来预测一个单词在特定话语中最可能的发音。pi先前的研究已经证实,说话者在各种情况下更有可能删除某些声音(例如,“about”中的最后一个“t”):如果下一个单词以某些辅音开头;如果他们说得很快;如果这个词是不令人惊讶的(有一个高的三重概率);如果他们是年轻的男性;如果单词后面没有停顿或单词重复;或者如果他们说某种方言。他们进一步展示了如何使用信心指标来识别基于语音实际发音的不准确的字典发音。这个项目将扩展和结合这两方面的研究,通过使用语音自信模型来识别导致单词错误的不正确发音的单词,然后为这些单词建立发音变化的统计(决策树)模型。pi的模型提供了一个通用引擎,他们将应用于几个特定类别的发音变化,包括方言和说话人类型,嵌入在Sphinx-II语音识别器中。将计算模型应用于语音变化的研究是语言学和认知建模领域的一个重要的新研究领域,对包括普遍获取在内的许多科学和教育学领域都有影响和潜在的好处。
英文摘要
In order for speech-based interfaces to be universally accessible - that is, available to all citizens in all situations - they must deal successfully with massive variability in pronunciation, as studies have shown that this is by far the single largest source of error in machine speech recognition. This project will address that issue by building dynamic pronunciation models which use context to predict the most likely pronunciations of a word in a particular utterance. The PIs' previous work has established that speakers are more likely to delete certain sounds (e.g., the final "t" in "about") in a variety of situations: if the next word starts with certain consonants; if they are speaking quickly; if the word is unsurprising (has a high trigram probability); if they are young and male; if the word is not followed by a pause or a word-repetition; or if they are speakers of certain dialects. They have further shown how confidence metrics can be used to identify inaccurate dictionary pronunciations based on actual pronunciations in speech. This project will extend and combine both of these lines of research, by using phonetic-confidence models to identify words with incorrect pronunciations that contribute to word-error, and then building statistical (decision-tree) models of pronunciation variation for these words. The PIs' model provides a general engine which they will apply to several specific classes of pronunciation variation, including dialect and speaker type, embedded in the Sphinx-II speech recognizer. The application of computational models to the study of pronunciation variation is an important new research area in linguistics and cognitive modeling, with implications for and potential benefits to many areas of science and pedagogy, including universal access.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: New tools for studying structural and inductive bias in NLP models
-
批准号:2128145
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2021
-
负责人:Daniel Jurafsky
-
依托单位:
RI: Medium: Deep Understanding: Integrating Neural and Symbolic Models of Meaning
-
批准号:1514268
-
项目类别:Continuing Grant
-
资助金额:$110.0万
-
财政年份:2015
-
负责人:Daniel Jurafsky
-
依托单位:
RI: Small: Learning Meaning and Grammar from Interaction, Context, and the World
-
批准号:1216875
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2012
-
负责人:Daniel Jurafsky
-
依托单位:
RI-Small: Unsupervised Learning of Meaning
-
批准号:0811974
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2008
-
负责人:Daniel Jurafsky
-
依托单位:
CAREER: Spoken Lexical Processing in Humans and Machines
-
批准号:9733067
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:1998
-
负责人:Daniel Jurafsky
-
依托单位:
SGER: Using Text Coherence and Verbal Valence in Long- Distance N-grams
-
批准号:9704046
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:1997
-
负责人:Daniel Jurafsky
-
依托单位:
海外基金