SGER: Self-Supervised Discriminative Training of Statistical Language Models
SGER: Self-Supervised Discriminative Training of Statistical Language Models
批准号:
0840112
负责人:
Sanjeev Khudanpur
金额:
$0.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2010-02-28
中文摘要
这一小笔探索性研究基金正在研究用于各种人类语言技术的统计语言模型的鉴别训练的新方法,如自动语音识别(ASR)和机器翻译(MT)。语言模型(LM)通常是通过正则化的最大似然法从目标领域的大量文本语料库中估计出来的。鉴别标准在ASR中得到了一些成功的应用,但由于需要额外的转录语音语料库来区分正确的单词序列和不正确的队列,这些标准的巨大前景受到了限制。这个项目正在探索在不需要大量人工标注的情况下区分估计语言模型的方法,即ASR的转录语音或MT的平行文本。正在探索的关键思想是,如果有大量(比方说)单语言中文文本,那么通过尝试使用现有的机器翻译系统将该文本翻译成(比方说)英语并检查哪些英语单词和短语最频繁地相互竞争,可以准确地估计汉语单词和短语的机器翻译队列。在任何特定的情况下,没有必要知道队列集合中的哪些相互竞争的单词或短语是正确的翻译!只需了解谁最常参加竞争就够了。研究人员正在使用单一语言的英语文本来探索区分队列集合中每个成员及其假定竞争对手的观察到的发病率的特征;从而综合得出区分训练的数据。他们正在调查这种有区别的训练是否专门针对机器翻译系统所面临的最令人衰弱的歧义。这个项目通过探索统计语言模型,使ASR和MT研究社区都受益,这些统计语言模型可以在不需要人工干预的情况下适应不断变化的任务或语言使用,并且对人工标注数据的依赖程度较低。ASR和机器翻译的进步反过来又将促进以多种语文和媒体更有效地利用计算机辅助获取信息。
英文摘要
Title: Self-Supervised Discriminative Training of Statistical Language ModelsThis Small Grant for Exploratory Research is investigating novel methods for discriminative training of statistical language models for application to various human language technologies, such as automatic speech recognition (ASR) and machine translation (MT).A language model (LM) is conventionally estimated from a large corpus of text in the target domain via regularized maximum likelihood. Discriminative criteria have been used with some success in ASR, but their immense promise has been curtailed by the requirement of an additional corpus of transcribed speech needed to discriminate between correct word sequences and their incorrect ?cohorts.? This project is exploring ways to discriminatively estimate language models without requiring massive manual annotation, namely, transcribed speech for ASR or parallel text for MT.The key idea being explored is that if a large amount of (say) monolingual Chinese text is available, then the MT cohorts of Chinese words and phrases may be accurately estimated by attempting to translate this text into (say) English using an existing MT system and examining which English words and phrases are most frequently in competition with each other. It is not necessary to know which of the competing words or phrases in a cohort set is the correct translation in any particular instance! It suffices to learn who are most often in competition. The investigators are using monolingual English text to explore features that discriminate between observed incidences of each member of a cohort set and its putative competitors; the data for discriminative training are thus derived synthetically. They are investigating if such a discriminatively trained LM specifically targets the most debilitating ambiguities faced by the MT system. The ASR counterpart, with cohort sets derived from automatic transcription of unannotated speech, is also being explored.This project benefits both the ASR and MT research communities by exploring statistical language models that can adapt without human intervention to changing tasks or language-use, and that are less reliant on manually annotated data. Advances in ASR and MT in turn will facilitate more effective computer-aided access to information in multiple languages and media.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CCRI: ENS: Next Generation Tools for Spoken Language Science & Technology
-
批准号:2120435
-
项目类别:Standard Grant
-
资助金额:$184.0万
-
财政年份:2021
-
负责人:Sanjeev Khudanpur
-
依托单位:
RI: Medium: Collaborative Research: Semi-Supervised Discriminative Training of Language Models
-
批准号:0963898
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2010
-
负责人:Sanjeev Khudanpur
-
依托单位:
Cross-Cutting Research Workshops on Intelligent Information Systems
-
批准号:1005411
-
项目类别:Continuing Grant
-
资助金额:$43.97万
-
财政年份:2010
-
负责人:Sanjeev Khudanpur
-
依托单位:
PIRE: Investigation of Meaning Representations in Language Understanding for Machine Translation Systems
-
批准号:0530118
-
项目类别:Continuing Grant
-
资助金额:$249.84万
-
财政年份:2005
-
负责人:Sanjeev Khudanpur
-
依托单位:
SGER: Pronunciation Modeling for Conversational Speech Recognition
-
批准号:9714169
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:1997
-
负责人:Sanjeev Khudanpur
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Self-DNA介导的CD4+组织驻留记忆T细胞(Trm)分化异常在狼疮肾炎发病中的作用及机制研究
-
批准号:82371813
-
项目类别:面上项目
-
资助金额:50万元
-
批准年份:2023
-
负责人:熊思东
-
依托单位:
基于受体识别和转运整合的self-DNA诱导采后桃果实抗病反应的机理研究
-
批准号:32302161
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:黎春红
-
依托单位:
基于广义测量的多体量子态self-test的实验研究
-
批准号:12104186
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:边志浩
-
依托单位:
Self-shrinkers的刚性及相关问题
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2019
-
负责人:魏国新
-
依托单位:
基于Self-peptide和Fe5C2构建的高敏感MR分子探针对肿瘤血管的MR靶向成像研究
-
批准号:81501521
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2015
-
负责人:龚明福
-
依托单位:
平均曲率流中非紧Self-shrinkers的结构
-
批准号:11301190
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2013
-
负责人:张坤
-
依托单位:
2维伪欧氏空间下平均曲率流中Self-shrinker问题的研究
-
批准号:11126152
-
项目类别:数学天元基金项目
-
资助金额:3.0万元
-
批准年份:2011
-
负责人:刘华侨
-
依托单位:
晶态桥联聚倍半硅氧烷的自导向组装(self-directed assembly)及其发光性能
-
批准号:21171046
-
项目类别:面上项目
-
资助金额:55.0万元
-
批准年份:2011
-
负责人:李焕荣
-
依托单位:
成束蛋白Fascin1在肺癌"self-seeding"过程中的作用及机制研究
-
批准号:81001041
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2010
-
负责人:赵晋波
-
依托单位:
工业用腈水合酶全新蛋白质翻译后调节体系self-subunit swapping的研究
-
批准号:31070711
-
项目类别:面上项目
-
资助金额:35.0万元
-
批准年份:2010
-
负责人:周哲敏
-
依托单位: