RI: Medium: Collaborative Research: Semi-Supervised Discriminative Training of Language Models
RI: Medium: Collaborative Research: Semi-Supervised Discriminative Training of Language Models
批准号:
0963898
负责人:
Sanjeev Khudanpur
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-06-01 至 2015-08-31
中文摘要
该项目正在进行统计语言建模的基础研究,以改善人类语言技术,包括自动语音识别(ASR)和机器翻译(MT)。语言模型(LM)通常使用目标语言的文本进行优化,以将高概率分配给格式良好的句子。这种方法有一个根本的缺点:优化没有明确地针对完成手头任务所需的各种区别,例如区分(对于ASR)不同的声音混淆的单词或(对于MT)表达多义源语言单词的多种含义的不同目标语言单词。LM的区分性优化将克服这一缺点,需要大量成对的输入-输出序列:语音及其用于ASR或源语言(例如中文)句子的参考转录,以及它们到MT的目标语言(例如英语)的翻译。这些资源是昂贵的,并且限制了区别性训练方法的功效。在一个彻底背离惯例,该项目正在研究歧视性训练使用容易获得的,* 未配对 * 输入和输出序列:未转录的语音或单语源语言文本和未配对的目标语言文本。正在追求两个关键思想:(i)处理未标记的输入序列(例如语音或中文文本)以学习ASR或MT系统可能遇到的混淆;(ii)利用未配对的输出序列(英语文本)来区分这些格式良好的句子与系统可能混淆的(假定的)格式不良的句子。这种自我监督的区别训练如果成功,将从根本上提高机器智能,影响许多其他应用。
英文摘要
This project is conducting fundamental research in statistical language modeling to improve human language technologies, including automatic speech recognition (ASR) and machine translation (MT). A language model (LM) is conventionally optimized, using text in the target language, to assign high probability to well-formed sentences. This method has a fundamental shortcoming: the optimization does not explicitly target the kinds of distinctions necessary to accomplish the task at hand, such as discriminating (for ASR) between different words that are acoustically confusable or (for MT) between different target-language words that express the multiple meanings of a polysemous source-language word. Discriminative optimization of the LM, which would overcome this shortcoming, requires large quantities of paired input-output sequences: speech and its reference transcription for ASR or source-language (e.g. Chinese) sentences and their translations into the target language (say, English) for MT. Such resources are expensive, and limit the efficacy of discriminative training methods. In a radical departure from convention, this project is investigating discriminative training using easily available, *unpaired* input and output sequences: un-transcribed speech or monolingual source-language text and unpaired target-language text. Two key ideas are being pursued: (i) unlabeled input sequences (e.g. speech or Chinese text) are processed to learn likely confusions encountered by the ASR or MT system; (ii) unpaired output sequences (English text) are leveraged to discriminate between these well-formed sentences from the (supposed) ill-formed sentences the system could potentially confuse them with. This self-supervised discriminative training, if successful, will advance machine intelligence in fundamental ways that impact many other applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CCRI: ENS: Next Generation Tools for Spoken Language Science & Technology
-
批准号:2120435
-
项目类别:Standard Grant
-
资助金额:$184.0万
-
财政年份:2021
-
负责人:Sanjeev Khudanpur
-
依托单位:
Cross-Cutting Research Workshops on Intelligent Information Systems
-
批准号:1005411
-
项目类别:Continuing Grant
-
资助金额:$43.97万
-
财政年份:2010
-
负责人:Sanjeev Khudanpur
-
依托单位:
SGER: Self-Supervised Discriminative Training of Statistical Language Models
-
批准号:0840112
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2008
-
负责人:Sanjeev Khudanpur
-
依托单位:
PIRE: Investigation of Meaning Representations in Language Understanding for Machine Translation Systems
-
批准号:0530118
-
项目类别:Continuing Grant
-
资助金额:$249.84万
-
财政年份:2005
-
负责人:Sanjeev Khudanpur
-
依托单位:
SGER: Pronunciation Modeling for Conversational Speech Recognition
-
批准号:9714169
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:1997
-
负责人:Sanjeev Khudanpur
-
依托单位:
海外基金