课题基金 / 基金详情

Collaborative Research: CompCog: Psychological, Computational, and Neural Adequacy in a Deep Learning Model of Human Speech Recognition

Collaborative Research: CompCog: Psychological, Computational, and Neural Adequacy in a Deep Learning Model of Human Speech Recognition
合作研究:CompCog:人类语音识别深度学习模型中的心理、计算和神经充分性
批准号:
2043950
负责人:
Kevin Brown
金额:
$17.89万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-06-01 至 2025-05-31

项目摘要

项目成果

Kevin Brown的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年里,语音识别的计算机技术已经取得了惊人的进步。我们中的许多人每天都使用它——在智能手机上口述短信或在自动电话系统中导航。尽管这些系统很好,但在复杂、拥挤和嘈杂的声学环境中,人类的表现仍然优于它们。如果对人类如何适应这些具有挑战性的情况有更多的了解,语音技术可能会变得更具适应性和鲁棒性。例如,用于语音识别的计算机系统使用复杂的“深度学习”网络,这些网络通常需要以与人类学习语言的方式截然不同的方式进行训练。尽管旨在模拟人类语言处理的神经网络模型要简单得多,这使得科学家能够对人类语言处理的工作原理进行假设,但它们并不使用真实的语音作为输入。相反,它们使用的语音特征更像是文本,而不是语音,因此无法解决人类如何将语音的声学映射到单词的核心问题。本研究项目的重点是弥合当前语音识别技术中使用的复杂人工神经网络模型与用于研究人类如何实际感知语音的简单神经网络模型之间的差距。本研究项目建立在一个新的语音神经网络模型上,该模型旨在对多个说话者产生的多个单词实现高识别精度。至关重要的是,该模型可以以最小的复杂性(使用比商业语音识别系统少得多的层)做到这一点,这使得研究人员能够理解它所执行的计算。研究计划包括将该模型扩展到更大的词汇量,对自然语言进行训练,并以人类听觉路径为模型添加生物学上合理的预处理。该模型将与人类口语单词识别行为的关键方面以及人类对口语的神经反应进行比较。这项工作有可能产生新的见解,通过使语音技术在具有挑战性的环境中更加强大,从而推动语音技术的发展,并对用于卫生、法律、教育的语音技术以及使聋人和重听人能够使用语音的自动字幕产生潜在影响。此外,从高中生到博士生都将成为研究团队的一员,他们将拥有丰富的研究经验,这将促进对学术研究或各种非学术职业有用的技术技能的发展。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Computer technology for speech recognition has advanced to an amazing degree over the past decade. Many of us use it daily -- to dictate text messages on smart phones or to navigate automated phone systems. As good as these systems are, humans still outperform them in complex, crowded, and noisy acoustic environments. If more were known concerning how humans adapt to these challenging situations, speech technology might be made more adaptive and robust. For example, computer systems for speech recognition use complex "deep learning" networks that often need to be trained in ways that are very different from how humans learn language. Although neural network models aimed at simulating human language processing are much simpler, which allows scientists to develop hypotheses about how human language processing works, they don't use real speech as input. Instead, they use phonetic features that are more like text than speech and so fail to address the core problem of how humans map the acoustics of speech to words. This research program focuses on bridging the gap between the complex artificial neural network models used in current technologies for speech recognition and the simpler neural network models used to investigate how humans actually perceive speech.This research program builds on a new neural network model for speech that aims to achieve high recognition accuracy on many words produced by several speakers. Crucially, the model can do this with minimal complexity (using many fewer layers than commercial speech recognition systems), which allows researchers to understand the computations it performs. The research plans include extending the model to a large vocabulary, training on naturalistic speech, and adding biologically plausible preprocessing modeled on the human auditory pathways. The model will be compared with key aspects of human spoken word recognition behavior as well as with human neural responses to spoken speech. The work has the potential to generate new insights to advance speech technology by making it more robust in challenging environments, with potential impact on speech technology used for health, law, education, and the automatic captioning that makes speech accessible to the deaf and hard of hearing. In addition, individuals ranging from high school students to Ph.D. students will be part of the research team and will have rich research experiences that will promote development of technical skills useful for careers in academic research or a variety of non-academic careers.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.3389/frai.2023.1062230
发表时间: 2023
期刊: FRONTIERS IN ARTIFICIAL INTELLIGENCE
影响因子: 4
作者: [Avcu, Enes, Hwang, Michael, Brown, Kevin Scott, Gow, David W.]
通讯作者: Gow, David W.
DOI: 10.1111/cogs.13291
发表时间: 2023-05-01
期刊: COGNITIVE SCIENCE
影响因子: 2.5
作者: [Brown,Kevin S., Yee,Eiling, McRae,Ken]
通讯作者: McRae,Ken
HSI Planning Project: Examining Inclusion and Other Variables of STEM Retention at a Faith-Based, Residential University
  • 批准号:
    2345328
  • 项目类别:
    Standard Grant
  • 资助金额:
    $19.52万
  • 财政年份:
    2024
  • 负责人:
    Kevin Brown
  • 依托单位:
CRCNS US-Spain Research Proposal: Collaborative Research: Tracking and modeling the neurobiology of multilingual speech recognition
  • 批准号:
    2207747
  • 项目类别:
    Standard Grant
  • 资助金额:
    $18.65万
  • 财政年份:
    2022
  • 负责人:
    Kevin Brown
  • 依托单位:
Collaborative Research: Modeling Spatiotemporal Control of EGFR-ERK Signaling in Gene-edited Cell Systems
  • 批准号:
    1906161
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2018
  • 负责人:
    Kevin Brown
  • 依托单位:
Collaborative Research: Modeling Spatiotemporal Control of EGFR-ERK Signaling in Gene-edited Cell Systems
  • 批准号:
    1715342
  • 项目类别:
    Standard Grant
  • 资助金额:
    $16.47万
  • 财政年份:
    2017
  • 负责人:
    Kevin Brown
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)