课题基金 / 基金详情

Collaborative Research: NCS-FO: Studying language in the brain in the modern machine learning era

Collaborative Research: NCS-FO: Studying language in the brain in the modern machine learning era
合作研究:NCS-FO:研究现代机器学习时代大脑中的语言
批准号:
2123818
负责人:
Gabriel Krieman
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-15 至 2025-08-31

项目摘要

项目成果

Gabriel Krieman的其他基金

相似基金

相关文献

中文摘要
翻译
该项目将研究大脑如何处理语言,这是我们能提出的最重要的问题之一。语言技能会显著影响一生的收入和社会差距。语言的丧失或语言发展的停滞可能是毁灭性的。与此同时,来自大脑的见解可以提高机器对语言的理解,开辟新的应用——从网络搜索到语音助手,有朝一日,机器人可以在我们的日常生活中帮助我们。用神经科学来理解语言使用过程中发生了什么,什么是对的,什么是错的,大脑使用什么语言结构和理论,将是革命性的。为了做到这一点,神经科学家使用了许多与机器学习相同的工具。这些机器学习工具在使用大型数据集方面有了巨大的改进,改变了机器的能力;然而,语言神经科学在很大程度上无法获得这些成果。我们将提供数据、新方法和指标,使神经科学能够扩大规模,并利用现代机器学习。与此同时,机器学习的规模使工具的使用民主化;科学界可以调查与他们有关的问题。今天,只有少数几个团体有资源来收集数据和调查语言神经科学方面的问题,使许多社区处于黑暗之中。一个大规模的数据、工具和基准的中央存储库将使大脑语言研究民主化,这是使我们成为人类的核心方面之一。我们的技术目标是产生最大的数据集,以1000倍的速度,用于研究语言的神经科学,以及利用这些数据的新型模型,以及形式化语言问题的基准,以获得关于语言网络和语言结构的见解。到目前为止,语言神经科学的研究只能在不同的数据集上提供语言网络的小快照,这使得很难建立一个连贯的画面。一个具有精确基准的单一大规模数据集,可以正式定义语言学中的假设在神经数据方面的含义,这将使社区能够对相同的数据提出许多问题,从而允许语言网络的结构和操作的综合。同时,已知需要大规模数据来探索对人工语言模型的理解。很可能,如果需要数万个句子来探测人工语言神经网络并获得有意义的见解,那么探测生物语言神经网络将需要相同规模的每个主题的数据。在公共数据集上围绕基准问题的形式化导致了从解析(Penn Treebank)到图像识别(ImageNet)等许多领域的天文数字进展;我们将把同样的方法应用于语言神经科学。这个过程非常有效,部分原因是它以非领域专家可以访问的方式提出问题;机器学习专家不需要关心语言的细节,他们将能够改进大脑对语言的解码,从而通过遵循现有协议来推动洞察力。通过以精确定义的方式提出语言问题,我们还将实现跨学科合作:语言学家将能够提出基准,这些基准是围绕分类器的性能或网络与神经活动之间的映射的问题。这些基准测试将提供一种通用的数学语言,不同的领域可以通过这种语言来表达他们的关键问题,这在以前是不可能的,因为没有数据集可以支持这样的工作。我们看到一个未来,神经科学、语言学、自然语言处理和机器学习作为一个整体,提出关于大脑语言的正确问题,开发支持回答这些问题的新工具,并探索支持构建语言系统连贯图像的大规模资源。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The project will investigate how the brain processes language, one of the most consequential questions we can ask. Language skills significantly affect lifetime income and social disparities. The loss of language or a halt in its development can be devastating. At the same time, insights from the brain that could improve machines’ understanding of language opens up new applications -- from web search, to voice assistants, to, one day, robots that can help us in our daily lives. Using neuroscience to understand what happens during language use, what goes right and wrong, what linguistic structures and theories are used by the brain, would be revolutionary. To do this, neuroscientists use many of the same tools as those created for machine learning. Those machine learning tools have improved tremendously using large datasets, changing what machines are capable of; yet the neuroscience of language has been largely unable to reap these rewards. We will provide that data, the new methods and metrics, required to enable neuroscience to scale up and take advantage of modern machine learning. At the same time, scale in machine learning has democratized access to tools; scientific communities can investigate questions that pertain to them. Today, only a few groups have the resources to collect data and investigate questions around the neuroscience of language, leaving many communities in the dark. A large-scale central repository of data, tools, and benchmarks will democratize access to the study of language in the brain, one of the core aspects of what makes us human.Our technical goal is to produce the largest dataset, by a factor of 1000, for investigating the neuroscience of language along with new types of models that exploit this data, and benchmarks which formalize linguistic questions to derive insights about the language network and the structure of language. Thus far, investigations in the neuroscience of language have only been able to provide small snapshots of the language network on different datasets, making it hard to build a coherent picture. A single large-scale dataset with precise benchmarks that formally define what hypotheses in linguistics mean in terms of neural data will enable the community to ask many questions of the same data, allowing for a synthesis of the structure and operation of the language network. At the same time, large-scale data is known to be required to probe the understanding of artificial language models. It is likely that if tens of thousands of sentences are required to probe an artificial language-neural-networks and derive meaningful insight, the same scale of data will be required per subject to probe biological language-neural-networks. Formalizing questions around benchmarks on a common dataset has resulted in astronomical progress in many fields from parsing (Penn Treebank) to image recognition (ImageNet); we will apply this same methodology to the neuroscience of language. This process is so efficient, in part, because it casts questions in a way that non-domain experts can access; machine learning experts need not concern themselves with linguistic minutia, they will be able to improve decoding of language from the brain and thereby drive insights by following existing protocols. By putting forward linguistic questions in a precisely-defined manner, we will also enable cross-disciplinary collaboration: linguists will be able to propose benchmarks which are questions around the performance of classifiers or mapping between networks and neural activity. These benchmarks will provide a common mathematical language by which different fields can express their key questions, in a way that has not been possible before because no dataset existed that could even support such work. We see a future where neuroscience, linguistics, natural language processing, and machine learning act as an integrated whole to ask the right questions about language in the brain, to develop new tools that support answering those questions, and to probe a large-scale resource that supports building a coherent picture of the language system.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Top-down processes to extract meaning from images
  • 批准号:
    1745365
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.0万
  • 财政年份:
    2017
  • 负责人:
    Gabriel Krieman
  • 依托单位:
Neurophysiological circuits underlying episodic memory formation in the human brain
  • 批准号:
    1358839
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $69.59万
  • 财政年份:
    2014
  • 负责人:
    Gabriel Krieman
  • 依托单位:
US-German Collaboration: Integration of Bottom-Up and Top-Down Signals in Visual Recognition
  • 批准号:
    1010109
  • 项目类别:
    Standard Grant
  • 资助金额:
    $30.6万
  • 财政年份:
    2010
  • 负责人:
    Gabriel Krieman
  • 依托单位:
CAREER:Deciphering the Neural Code From Perception To Cognition
  • 批准号:
    0954570
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $50.34万
  • 财政年份:
    2010
  • 负责人:
    Gabriel Krieman
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)