课题基金 / 基金详情

MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling

MRI: Acquisition of the LanguageLens for Large-Scale Language Modeling
MRI:获取用于大规模语言建模的 LanguageLens
批准号:
2214708
负责人:
David Wingate
金额:
$101.48万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2025-07-31

项目摘要

项目成果

David Wingate的其他基金

相似基金

相关文献

中文摘要
翻译
机器学习正在彻底改变社会的许多方面,但训练最好的模型需要大量的计算资源,而这些资源往往是学术团体无法企及的。 因此,该项目获得了一种名为“人工智能透镜”的专用仪器,旨在处理大量的自然语言文本。 该平台将支持自然语言处理、深度学习、计算语言学、危机信息学、会话人工智能、神经机器翻译和法律的语料库语言学等领域的研究,并将使学术研究能够推进训练大型模型所需的机器学习以及这些模型的社会相关应用。该平台是一个高性能的GPU集群,可以平衡计算、存储和节点间通信,以支持各种要求苛刻的基于NLP的工作负载。该项目将专注于解决那些有潜力在各个领域产生变革性、跨学科影响的研究项目。 该工具的一个重点领域是能够训练新的大规模语言模型并实时检查其内部工作。 语言模型将在新的语料库以及新的神经符号架构上进行特定下游应用的训练,以帮助从所得权重中获得洞察力。该基金将优先支持解决紧迫社会问题的研究。 它还将为学生提供真实的劳动力培训和教育体验:随着工业界和学术界之间的资源差距越来越大,越来越难以为他们提供机会进行涉及庞大模型和数据集的高影响力研究。 最后,由于许多公司拒绝发布其模型的预训练权重,因此一个核心目标是在道德考虑的前提下,向所有人免费提供训练权重,以推动行业和学术界的全国影响。项目资源,如代码,出版物,数据集和预先训练的模型将通过www.example.com网站提供https://ll.cs.byu.edu/.This奖项反映了NSF的法定使命,并被认为值得通过使用基金会的知识价值和更广泛的影响审查标准进行评估来支持。
英文摘要
Machine learning is revolutionizing many parts of society, but training the very best models requires tremendous computing resources that are often out of reach for academic groups. This project therefore acquires a special-purpose instrument, named the LanguageLens, that is designed to process vast amounts of natural language text. The LanguageLens will support research in natural language processing, deep learning, computational linguistics, crisis informatics, conversational AI, neural machine translation, and legal corpus linguistics, and will enable academic research to advance both the machine learning needed to train large models, as well as societially relevant applications of those models.The LanguageLens is a high-performance GPU cluster that balances compute, storage and internode communication to support a variety of demanding NLP-based workloads. The LanguageLens will be focused on solving research projects that have the potential for transformational, interdisciplinary impact across a wide variety of fields. A key area of focus for the instrument is the ability to train new large-scale language models and to examine their inner workings in real-time. Language models will be trained with specific downstream applications in mind, on novel corpora as well as with novel neuro-symbolic architectures, to help derive insight from the resulting weights. The LanguageLens will prioritize support for research that addresses pressing societal problems. It will also provide authentic workforce training and educational experiences for students: as the resource gap between industry and academia grows, it is increasingly difficult to give them opportunities to pursue high-impact research that involves huge models and datasets. Finally, as many companies refuse to release the pretrained weights of their models, a central goal is to make trained weights freely available to everyone, subject to ethical considerations, to drive national impact for both industry and academia. Project resources such as code, publications, datasets and pretrained models will be available through the LanguageLens website at https://ll.cs.byu.edu/.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Harnessing Accurate Bias in Large-Scale Language Models
  • 批准号:
    2141680
  • 项目类别:
    Standard Grant
  • 资助金额:
    $27.89万
  • 财政年份:
    2021
  • 负责人:
    David Wingate
  • 依托单位:
CAREER: Blending Deep Reinforcement Learning and Probabilistic Programming
  • 批准号:
    1652950
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.97万
  • 财政年份:
    2017
  • 负责人:
    David Wingate
  • 依托单位:
海外基金