课题基金 / 基金详情

MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling

MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling
MRI:获取用于大规模语言建模的 LanguageLens
批准号:
2214708
负责人:
David Wingate
金额:
$101.48万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2025-07-31

项目摘要

项目成果

David Wingate的其他基金

相似基金

相关文献

中文摘要
翻译
机器学习正在彻底改变社会的许多方面,但训练最好的模型需要大量的计算资源,这往往是学术团体无法企及的。因此,这个项目需要一个特殊用途的工具,名为LanguageLens,用于处理大量的自然语言文本。LanguageLens将支持自然语言处理、深度学习、计算语言学、危机信息学、会话人工智能、神经机器翻译和法律语料库语言学等领域的研究,并将使学术研究能够推进训练大型模型所需的机器学习,以及这些模型的社会相关应用。LanguageLens是一个高性能GPU集群,可以平衡计算、存储和节点间通信,以支持各种要求苛刻的基于nlp的工作负载。LanguageLens将专注于解决在各个领域具有变革性、跨学科影响的潜在研究项目。该仪器的一个重点领域是能够训练新的大规模语言模型并实时检查其内部工作原理。语言模型将在特定的下游应用程序中进行训练,在新的语料库以及新的神经符号体系结构上,以帮助从结果权重中获得洞察力。语言elens将优先支持解决紧迫社会问题的研究。它还将为学生提供真实的劳动力培训和教育经验:随着产业界和学术界之间资源差距的扩大,越来越难以给他们机会从事涉及庞大模型和数据集的高影响力研究。最后,由于许多公司拒绝发布其模型的预训练权重,一个中心目标是在道德考虑的前提下,让每个人都可以免费获得训练过的权重,从而推动对行业和学术界的国家影响。项目资源,如代码、出版物、数据集和预训练模型将通过LanguageLens网站https://ll.cs.byu.edu/.This提供。该奖项反映了国家科学基金会的法定使命,并通过基金会的智力价值和更广泛的影响审查标准进行评估,认为值得支持。
英文摘要
Machine learning is revolutionizing many parts of society, but training the very best models requires tremendous computing resources that are often out of reach for academic groups. This project therefore acquires a special-purpose instrument, named the LanguageLens, that is designed to process vast amounts of natural language text. The LanguageLens will support research in natural language processing, deep learning, computational linguistics, crisis informatics, conversational AI, neural machine translation, and legal corpus linguistics, and will enable academic research to advance both the machine learning needed to train large models, as well as societially relevant applications of those models.The LanguageLens is a high-performance GPU cluster that balances compute, storage and internode communication to support a variety of demanding NLP-based workloads. The LanguageLens will be focused on solving research projects that have the potential for transformational, interdisciplinary impact across a wide variety of fields. A key area of focus for the instrument is the ability to train new large-scale language models and to examine their inner workings in real-time. Language models will be trained with specific downstream applications in mind, on novel corpora as well as with novel neuro-symbolic architectures, to help derive insight from the resulting weights. The LanguageLens will prioritize support for research that addresses pressing societal problems. It will also provide authentic workforce training and educational experiences for students: as the resource gap between industry and academia grows, it is increasingly difficult to give them opportunities to pursue high-impact research that involves huge models and datasets. Finally, as many companies refuse to release the pretrained weights of their models, a central goal is to make trained weights freely available to everyone, subject to ethical considerations, to drive national impact for both industry and academia. Project resources such as code, publications, datasets and pretrained models will be available through the LanguageLens website at https://ll.cs.byu.edu/.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Harnessing Accurate Bias in Large-Scale Language Models
  • 批准号:
    2141680
  • 项目类别:
    Standard Grant
  • 资助金额:
    $27.89万
  • 财政年份:
    2021
  • 负责人:
    David Wingate
  • 依托单位:
CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
  • 批准号:
    1652950
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.97万
  • 财政年份:
    2017
  • 负责人:
    David Wingate
  • 依托单位:
海外基金