CAREER: Insertion-Based Natural Language Generation
CAREER: Insertion-Based Natural Language Generation
批准号:
2339766
负责人:
Nanyun Peng
金额:
$58.55万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-05-01 至 2029-04-30
中文摘要
语言模型已经成为当今大多数自然语言处理(NLP)应用的基础。然而,现有的语言模型主要遵循自回归范式,即训练模型在给定左侧上下文的情况下预测下一个单词。然后,他们从左到右逐字生成句子。这种范式的简单性很吸引人,但它也有几个局限性,包括生成效率低下,缺乏对人机协作的可靠控制,更重要的是,它显著偏离了人类解释和撰写句子的方式。这项拟议的研究探索了一种截然不同的语言建模范式-基于插入的模型。这种模型将生成过程描述为在不完整的上下文中迭代地插入单词,与自回归模型相比,提供了显著更大的灵活性和可控性。基于插入的公式更好地模拟了人类的写作行为,从而为计算语言学和认知科学提供了研究语言结构的工具。该模型的可控性将使通信研究人员、内容创建者和创造性作曲家受益于创造性内容生成、定制通信和个性化用户体验等应用。研究成果将被整合到教材中,为K-12和本科生传播生成性人工智能(AI)。本项目旨在促进我们对基于插入的LMS的好处和能力的理解。这项拟议的研究将探索基于插入的公式的灵活性所支持的不同的生成顺序,目的是在语言学理论的指导下找到最优的插入顺序。此外,该项目将为基于插入的LMS的变体引入一个新的模型体系结构,其中包括删除操作,使这些模型能够纠正上一代错误。这种增强为基于插入的LMS带来了更多的灵活性和可控性。此外,本研究的一个雄心勃勃的目标是研究基于插入的LMS的标度规律并扩大其预训练。这些探索的成功有可能导致一系列大型语言模型,这些模型表现出增强的灵活性、可控性和推理效率,超过自回归LMS,从而为广泛的自然语言生成和一般NLP任务带来好处。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
AbstractLanguage models (LMs) have become the foundations of most natural language processing (NLP) applications nowadays. However, existing language models predominantly follow the auto-regressive paradigm, which trains models to predict the next word given the left-side context. Then, they generate sentences word-by-word, left-to-right. The simplicity of this paradigm is attractive but it has several limitations, including inefficiency in generation, lack of reliable control for human-machine collaboration, and more importantly, it significantly deviates from how humans interpret and compose sentences. The proposed research explores a fundamentally distinct paradigm of language modeling --- insertion-based models. Such models formulate the generation process as iteratively inserting words into an incomplete context, offering significantly more flexibility and controllability compared to the auto-regressive models. The insertion-based formulation better mimics human writing behaviors and thus provides a tool for computational linguistics and cognitive science to study the structure of languages. The controllability of the model will benefit communication researchers, content creators, and creative composers for applications such as creative content generation, tailored communication, and personalized user experiences. The research results will be integrated in teaching materials to disseminate generative artificial intelligence (AI) for K-12 and undergraduate. This project aims to advance our understandings of the benefits and capability of insertion-based LMs. The proposed research will explore different generation orders supported by the flexibility of the insertion-based formulation, with the goal to discover optimal insertion orders guided by linguistic theories. Additionally, the project will introduce a novel model architecture for a variation of insertion-based LMs that incorporates deletion operations, enabling the models to rectify previous generation errors. This enhancement brings added flexibility and controllability to insertion-based LMs. Furthermore, an ambitious goal of this research is to investigate the scaling law and scale up the pretraining of insertion-based LMs. The success of these explorations has the potential to lead to a family of large language models that exhibit enhanced flexibility, controllability, and inference efficiency surpassing the auto-regressive LMs, resulting in benefits for a wide range of natural language generation and general NLP tasks.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金