课题基金 / 基金详情

Improving language models, inspired by the brain

Improving language models, inspired by the brain
受大脑启发,改进语言模型
批准号:
RGPIN-2022-03580
负责人:
Fyshe, Alona
金额:
$2.99万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Fyshe, Alona的其他基金

相似基金

相关文献

中文摘要
翻译
语言模型(LMS)在过去10年中有了显著改进,这在很大程度上要归功于它们现在使用的深度学习神经网络模型。这些学习的模型可以被提供一个文本片段,并生成有意义地完成想法的单词。目前的模型可以生成非常像人类的文本,以至于很难将其与人写的文本区分开来。这些模型已经为《纽约客》撰写了文章片段,并可以准确地回答某些类型的阅读理解问题,准确度达到或超过普通人的准确度。事实上,这些模型非常准确,有时它们能够生成比人类生成的文本更好的新奇文本。值得注意的是,一个这样的模型生成了前一句话。然而,尽管目前的LMS很有说服力,但在某些情况下,它们仍然不完美,不是最优的。例如,当任务需要推理、概括或常识知识时,LMS的性能很差。它们还需要大量的数据来进行训练。相比之下,年幼的孩子只需要我们模型中的一小部分数据,他们是泛化大师和推理者。断线在哪里?为什么孩子们在语言方面如此精通,我们目前的模式根本无法复制?通过这个研究计划,我将研究语言开发和使用的模式,并将其作为改进计算LMS的灵感。我将研究婴儿的语言习得和语义组成的开始,儿童对不同类型信息的优先排序,以及流利的成年人控制语言使用的元认知过程。生命周期中的每一个阶段都集中了人类语言学习和使用的不同方面,以及理解和改进计算语言学习的新方法。这项工作的影响是多方面的。首先,这些发现将使我能够创建更准确但也更有效的语言模型。更高效的LMS将用更少的数据训练得更快。这将改善语言模型的多种下游使用,如对话生成和情感分类。因为更有效的语言模型也需要更少的资源来训练,我的工作将影响当前自然语言处理社区的碳足迹。其次,我提出的研究将有助于我们理解大脑是如何获得和发展语言的,反馈到学习的基本神经科学中,并有助于为扫盲教育研究提供信息。我的研究的影响将和研究项目本身一样是跨学科的。
英文摘要
Language models (LMs) have markedly improved over the past 10 years, due in large part to the deep learning neural networks models they now use. These learned models can be fed a text snippet and generate words to meaningfully finish the thought. Current models can generate text that is so human-like that it can be difficult to differentiate from text written by people. These models have written segments of articles for the New Yorker and can answer some types of reading comprehension questions with accuracy at or above the accuracy of the average person. In fact, these models are so accurate, they are sometimes able to generate novel text that is better than what a human could generate. Remarkably, one such model generated the previous sentence. However, as convincing as current LMs are, they remain imperfect and sub-optimal in certain scenarios. For example, LMs perform poorly when tasks require inference, generalization, or commonsense knowledge. They also require huge amounts of data to train. Contrast this to young children who require only a fraction of the data our models do, and who are master generalizers and inference-makers. Where is that disconnect? Why are children so masterful with language in ways that our current models simply cannot replicate? Through this proposed research plan, I will study patterns in language development and usage, and use them as inspiration for improving computational LMs. I will study language acquisition and the onset of semantic composition in infants, the prioritization of different kinds of information in children, and the meta-cognitive processes governing language usage in fluent adults. Each of these points in the lifespan bring into focus different aspects of human language learning and usage, and novel ways of understanding and improving computational language learning. The impact of this work is multi-faceted. Firstly, these findings will allow me to create more accurate but also more efficient language models. More efficient LMs will train faster, with less data. This will improve multiple downstream usages of language models like dialogue generation and sentiment classification. Because more efficient models of language also require fewer resources to train, my work will impact the carbon footprint of the current natural language processing community. Secondly, my proposed research will contribute to our understanding of how the brain acquires and develops language, feed back into the basic neuroscience of learning, and help inform literacy education research. The impacts of my research will be as interdisciplinary as the research program itself.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Computational Techniques for Exploring Language in the Brain
  • 批准号:
    RGPIN-2016-05265
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.6万
  • 财政年份:
    2021
  • 负责人:
    Fyshe, Alona
  • 依托单位:
Computational Techniques for Exploring Language in the Brain
  • 批准号:
    RGPIN-2016-05265
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.6万
  • 财政年份:
    2020
  • 负责人:
    Fyshe, Alona
  • 依托单位:
Computational Techniques for Exploring Language in the Brain
  • 批准号:
    RGPIN-2016-05265
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.6万
  • 财政年份:
    2019
  • 负责人:
    Fyshe, Alona
  • 依托单位:
Computational Techniques for Exploring Language in the Brain
  • 批准号:
    RGPIN-2016-05265
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.6万
  • 财政年份:
    2018
  • 负责人:
    Fyshe, Alona
  • 依托单位:
国内基金
海外基金
儿童音乐能力发展对语言与社会认知能力及脑发育的影响
  • 批准号:
    31971003
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    南云
  • 依托单位:
面向英汉双向跨语言图像检索的文本分析关键技术研究
  • 批准号:
    61170095
  • 项目类别:
    面上项目
  • 资助金额:
    57.0万元
  • 批准年份:
    2011
  • 负责人:
    张玥杰
  • 依托单位:
儿童植入耳蜗后听觉行为与言语发展进程的关联性研究
  • 批准号:
    81170916
  • 项目类别:
    面上项目
  • 资助金额:
    65.0万元
  • 批准年份:
    2011
  • 负责人:
    刘莎
  • 依托单位:
基于儿童心理分析的图解式汉语口语自动解析方法研究
  • 批准号:
    60175012
  • 项目类别:
    面上项目
  • 资助金额:
    18.0万元
  • 批准年份:
    2001
  • 负责人:
    宗成庆
  • 依托单位: