课题基金 / 基金详情

CAREER: Controllable generation for instruction-following language models

CAREER: Controllable generation for instruction-following language models
职业:指令跟随语言模型的可控生成
批准号:
2338866
负责人:
Tatsunori Hashimoto
金额:
$53.13万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-04-15 至 2029-03-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Instruction-following language models like ChatGPT are beginning to see widespread development, and the ability to understand these systems and control them is critically important to make sure that they benefit society. Despite the success of these language models in generating fluent and convincing-looking outputs, there has been a growing body of work indicating that these systems can generate outputs that are undesirable to users, model creators, and even society at large. This gap between the ability to create models that imitate humans and the inability to have them fulfill specific desiderata (e.g. refuse to generate incorrect information) shows a major deficiency in the ability to precisely control these systems. This project aims to build principled, transparent, and precise methods for controlling language models.To achieve these goals, this project views controllable generation as a viable long-term path to creating instruction-following language models that precisely follow our design goals. Controllable generation offers several benefits. First, it defines a precise statistical modeling problem on which it is possible to build principled methods and rigorous evaluations. Second, it separates the control target from the task, improving transparency by allowing users to see exactly what is being optimized by the model designers. Third, it enables much more precise controls via inference-time methods such as rejection sampling, which strictly enforces the control as a constraint. While controllable generation has major long-term benefits for language models, there also remain significant open problems that must be resolved first, including the difficulty of performing discrete search, the need for specialized training, and the lack of realistic benchmarks of control tasks in the wild. We will address these challenges through a combination of new models (such as diffusion-based models), zero-shot and decoder-based control methods, and a broad benchmark of in-the-wild control behaviors.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金