MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling
MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling
批准号:
2214708
负责人:
David Wingate
金额:
$101.48万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2025-07-31
中文摘要
机器学习正在给社会的许多方面带来革命性的变化,但训练最好的模型需要巨大的计算资源,而学术团体往往无法获得这些资源。因此,该项目获得了一种特殊用途的仪器,名为LanguageLens,旨在处理大量的自然语言文本。LanguageLens将支持自然语言处理、深度学习、计算语言学、危机信息学、对话人工智能、神经机器翻译和法律语料库语言学的研究,并将使学术研究能够促进训练大型模型所需的机器学习,以及这些模型的社交相关应用。LanguageLens是一个高性能的GPU集群,可平衡计算、存储和节点间通信,以支持各种基于NLP的苛刻工作负载。LanguageLens将专注于解决具有变革潜力的研究项目,这些项目将在广泛的领域产生跨学科影响。该工具的一个关键重点领域是训练新的大规模语言模型并实时检查其内部工作的能力。语言模型的训练将考虑到特定的下游应用,包括新的语料库以及新的神经符号体系结构,以帮助从结果权重中获得洞察力。LanguageLens将优先支持解决紧迫社会问题的研究。它还将为学生提供真实的劳动力培训和教育体验:随着产业界和学术界之间的资源差距扩大,让他们有机会从事涉及巨大模型和数据集的高影响力研究变得越来越困难。最后,由于许多公司拒绝公布其模型的预先训练的重量,一个核心目标是在符合道德考虑的情况下,向所有人免费提供训练过的重量,以推动行业和学术界在全国范围内的影响。代码、出版物、数据集和预先培训的模型等项目资源将通过https://ll.cs.byu.edu/.This的LanguageLens网站获得,该奖项反映了国家科学基金会的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine learning is revolutionizing many parts of society, but training the very best models requires tremendous computing resources that are often out of reach for academic groups. This project therefore acquires a special-purpose instrument, named the LanguageLens, that is designed to process vast amounts of natural language text. The LanguageLens will support research in natural language processing, deep learning, computational linguistics, crisis informatics, conversational AI, neural machine translation, and legal corpus linguistics, and will enable academic research to advance both the machine learning needed to train large models, as well as societially relevant applications of those models.The LanguageLens is a high-performance GPU cluster that balances compute, storage and internode communication to support a variety of demanding NLP-based workloads. The LanguageLens will be focused on solving research projects that have the potential for transformational, interdisciplinary impact across a wide variety of fields. A key area of focus for the instrument is the ability to train new large-scale language models and to examine their inner workings in real-time. Language models will be trained with specific downstream applications in mind, on novel corpora as well as with novel neuro-symbolic architectures, to help derive insight from the resulting weights. The LanguageLens will prioritize support for research that addresses pressing societal problems. It will also provide authentic workforce training and educational experiences for students: as the resource gap between industry and academia grows, it is increasingly difficult to give them opportunities to pursue high-impact research that involves huge models and datasets. Finally, as many companies refuse to release the pretrained weights of their models, a central goal is to make trained weights freely available to everyone, subject to ethical considerations, to drive national impact for both industry and academia. Project resources such as code, publications, datasets and pretrained models will be available through the LanguageLens website at https://ll.cs.byu.edu/.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Harnessing Accurate Bias in Large-Scale Language Models
-
批准号:2141680
-
项目类别:Standard Grant
-
资助金额:$27.89万
-
财政年份:2021
-
负责人:David Wingate
-
依托单位:
CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
-
批准号:1652950
-
项目类别:Continuing Grant
-
资助金额:$50.97万
-
财政年份:2017
-
负责人:David Wingate
-
依托单位:
海外基金