课题基金 / 基金详情

RI: Medium: Collaborative Research: Semi-Supervised Discriminative Training of Language Models

RI: Medium: Collaborative Research: Semi-Supervised Discriminative Training of Language Models
RI:媒介:协作研究:语言模型的半监督判别训练
批准号:
0963898
负责人:
Sanjeev Khudanpur
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-06-01 至 2015-08-31

项目摘要

项目成果

Sanjeev Khudanpur的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This project is conducting fundamental research in statistical language modeling to improve human language technologies, including automatic speech recognition (ASR) and machine translation (MT). A language model (LM) is conventionally optimized, using text in the target language, to assign high probability to well-formed sentences. This method has a fundamental shortcoming: the optimization does not explicitly target the kinds of distinctions necessary to accomplish the task at hand, such as discriminating (for ASR) between different words that are acoustically confusable or (for MT) between different target-language words that express the multiple meanings of a polysemous source-language word. Discriminative optimization of the LM, which would overcome this shortcoming, requires large quantities of paired input-output sequences: speech and its reference transcription for ASR or source-language (e.g. Chinese) sentences and their translations into the target language (say, English) for MT. Such resources are expensive, and limit the efficacy of discriminative training methods. In a radical departure from convention, this project is investigating discriminative training using easily available, *unpaired* input and output sequences: un-transcribed speech or monolingual source-language text and unpaired target-language text. Two key ideas are being pursued: (i) unlabeled input sequences (e.g. speech or Chinese text) are processed to learn likely confusions encountered by the ASR or MT system; (ii) unpaired output sequences (English text) are leveraged to discriminate between these well-formed sentences from the (supposed) ill-formed sentences the system could potentially confuse them with. This self-supervised discriminative training, if successful, will advance machine intelligence in fundamental ways that impact many other applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CCRI: ENS: Next Generation Tools for Spoken Language Science & Technology
  • 批准号:
    2120435
  • 项目类别:
    Standard Grant
  • 资助金额:
    $184.0万
  • 财政年份:
    2021
  • 负责人:
    Sanjeev Khudanpur
  • 依托单位:
Cross-Cutting Research Workshops on Intelligent Information Systems
  • 批准号:
    1005411
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $43.97万
  • 财政年份:
    2010
  • 负责人:
    Sanjeev Khudanpur
  • 依托单位:
SGER: Self-Supervised Discriminative Training of Statistical Language Models
  • 批准号:
    0840112
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2008
  • 负责人:
    Sanjeev Khudanpur
  • 依托单位:
PIRE: Investigation of Meaning Representations in Language Understanding for Machine Translation Systems
  • 批准号:
    0530118
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $249.84万
  • 财政年份:
    2005
  • 负责人:
    Sanjeev Khudanpur
  • 依托单位:
海外基金