SGER: Using Text Coherence and Verbal Valence in Long- Distance N-grams
SGER: Using Text Coherence and Verbal Valence in Long- Distance N-grams
批准号:
9704046
负责人:
Daniel Jurafsky
金额:
$5.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
1997
资助国家:
美国
项目状态:
已结题
起止时间:
1997-01-15 至 1997-12-31
中文摘要
***构建更好的语音识别器需要用复杂的概率语言知识来增强n-gram语法。这个项目是建立两个重要的句法/语义知识的概率模型:动词-论点约束和语义文本连贯。动词对其论点的语法和语义有很强的限制。这个项目是计算不同参数结构可能与不同动词同时出现的概率,并使用这些概率来增强标准的三元组语言模型。(2)语篇和话语在语义上趋于连贯;特别是文本中出现的单词往往在语义上彼此相关。该项目将一种称为潜在语义分析(LSA)的单词含义模型应用于ASR lm。在LSA中,词相似度度量是通过计算词共现概率的大矩阵来定义的,然后通过奇异值分解对其进行平滑处理,从而得到语义词相似度的广义度量。然后,三元组模型可以增加相似单词在彼此附近出现的概率。建立这两种语言知识的随机模型,除了可能应用于语音识别LMs、词义消歧或解析之外,还有助于弥合语言学中使用的结构模型与语音工程中使用的统计模型之间的差距
英文摘要
*** Building better speech recognizers requires augmenting n-gram grammars with sophisticated yet probabilistic linguistic knowledge. This project is building probabilistic models of two important pieces of syntactic/semantic knowledge: verb-argument constraints and semantic text coherence.(1) Verbs place strong constraints on the syntax and semantics of their arguments. This project is computing probabilities for the different argument structures that can co-occur with different verbs, and using these probabilities to augment standard trigram language models.(2) Texts and discourses tend to be semantically coherent;in particular the words that occur in a text tend to be semantically related to each other. This project is applying a model of word meaning called Latent Semantic Analysis (LSA) to ASR LMs. In LSA, a word-similarity metric is defined by computing a large matrix of word co-occurrence probabilities, which are then smoothed via Singular Value Decomposition, resulting in a generalized measure of semantic word-similarity. Trigram models can then increase the probability that similar words will occur near each other. Building these two stochastic models of linguistic knowledge, besides possible application in speech recognition LMs, word-sense disambiguation, or parsing, also helps bridge the gap between the structural models used in linguistics and the statistical models of speech engineering.***
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: New tools for studying structural and inductive bias in NLP models
-
批准号:2128145
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2021
-
负责人:Daniel Jurafsky
-
依托单位:
RI: Medium: Deep Understanding: Integrating Neural and Symbolic Models of Meaning
-
批准号:1514268
-
项目类别:Continuing Grant
-
资助金额:$110.0万
-
财政年份:2015
-
负责人:Daniel Jurafsky
-
依托单位:
RI: Small: Learning Meaning and Grammar from Interaction, Context, and the World
-
批准号:1216875
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2012
-
负责人:Daniel Jurafsky
-
依托单位:
RI-Small: Unsupervised Learning of Meaning
-
批准号:0811974
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2008
-
负责人:Daniel Jurafsky
-
依托单位:
Modeling Pronunciation Variation for Universal Access to Speech Understanding
-
批准号:9978025
-
项目类别:Continuing Grant
-
资助金额:$50.4万
-
财政年份:1999
-
负责人:Daniel Jurafsky
-
依托单位:
CAREER: Spoken Lexical Processing in Humans and Machines
-
批准号:9733067
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:1998
-
负责人:Daniel Jurafsky
-
依托单位:
国内基金
海外基金
Capture and Release of Droplets Using Advanced Materials for High Technology Applications
-
批准号:52073127
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2020
-
负责人:Alidad Amirfazli
-
依托单位:
Molecular Interaction Reconstruction of Rheumatoid Arthritis Therapies Using Clinical Data
-
批准号:31070748
-
项目类别:面上项目
-
资助金额:34.0万元
-
批准年份:2010
-
负责人:Christine Nardini
-
依托单位: