课题基金 / 基金详情

Enhancing research on speech and deep learning through holistic acoustic analysis

Enhancing research on speech and deep learning through holistic acoustic analysis
通过整体声学分析加强语音和深度学习研究
批准号:
2219843
负责人:
Matthew Goldrick
金额:
$100.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-15 至 2026-07-31

项目摘要

项目成果

Matthew Goldrick的其他基金

相似基金

相关文献

中文摘要
翻译
你可以从一个人的发音方式猜到很多关于他的事情。值得注意的是,人类听众可以判断说话者是否可能将英语作为第一语言或第二语言学习,或者说话者是否有大脑损伤,从而使他们难以说话。这种直觉依赖于人类听者的整体模式识别能力;这些能力使我们能够感知发音之间重要、有意义但又微妙的差异。然而,科学家目前用来客观测量语音的方法--基于语音的少量属性--未能捕捉到这些差异,阻碍了我们利用语音了解大脑和大脑的能力。这个项目将语音科学家、计算机科学家和神经学家聚集在一起,测试一种完全不同的方法来解决这个问题。机器学习将被用来发现一种基于整体模式识别的量化口语差异的新方法。这将根据来自双语使用者的新的和现有的数据进行测试。如果成功,这将产生一种完全通用的方法,可以应用于任何语言或任何语言用法的语音,使科学家能够利用语音中的丰富信息来开发对思维和大脑的强大新见解。改善对发音细微问题的检测,例如阿尔茨海默病,将促进我们对人类用来产生语音的大脑机制的理解。这项测试的结果还将使计算机科学家加深我们对机器学习算法如何处理声音的理解,推动算法的改进,并支持任何依赖口语处理的语音和语言技术领域的应用。说话者之间的语音差异为认知神经学家提供了宝贵的信息宝库,有助于深入了解语言处理的认知机制,并可能提供大脑功能障碍的早期迹象。当前对语音的研究由于需要预先选择特定的时间尺度和声学维度的分析而受阻。我们提出了一种完全不同的方法:使用无监督深度学习来发现声学变异分析的表征空间。为了测试这一高度通用的方法,我们将把这种方法与目前分析双语语音中个体差异的最先进方法进行比较。这包括利用第二语言语音中的声学变化来预测可理解性,并检测代码转换的困难,特别是阿尔茨海默病患者面临的挑战。这一结果将为深度学习和认知神经科学的发展提供信息。机器学习算法是完全通用的;它可以应用于任何语言或任何语言使用领域的语音,扩大了语音技术或认知神经科学家研究的人群和背景的范围。该项目的综合方法将使计算机科学家能够提高我们对现代深度学习架构在多大程度上接近人类语音处理的理解,并允许认知神经学家进一步理解有意义的声学区别是如何在语音感知和产生中表示的。人类语音表示法。本项目由理解神经和认知系统的综合策略(NCS)计划资助,该计划由计算机和信息科学与工程(CEISE)、教育和人力资源(EHR)、工程(ENG)和社会、行为和经济科学(SBE)等主管部门共同支持。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
You can guess a lot about a person from the way they pronounce words. Remarkably, human listeners can tell if it is likely that talkers learned English as a first language or a second language, or if the talkers might have a brain injury that makes it difficult for them to speak. Such intuitions rely on human listeners’ holistic pattern recognition abilities; these allow us to perceive the important, meaningful, yet subtle differences between pronunciations. However, the methods scientists currently use to measure speech objectively – based on a small number of properties of speech sounds – fail to capture these differences, hampering our ability to use speech to learn about the mind and brain. This project brings together speech scientists, computer scientists, and neuroscientists to test a radically different approach to this problem. Machine learning will be used to discover a new method for quantifying differences between spoken utterances based on holistic pattern recognition. This will be tested against new and existing data from bilingual speakers. If successful, this will yield a fully general method that can be applied to speech from any language or any domain of language usage, allowing scientists to capitalize on the wealth of information in speech to develop powerful new insights into the mind and brain. Improved detection of subtle problems with pronunciation, such as occurs with Alzheimer’s disease, will advance our understanding of the brain mechanisms that humans use to produce speech. The results of this testing will also allow computer scientists to advance our understanding of how machine learning algorithms process sounds, driving improvements in the algorithms and supporting applications in any area of speech and language technology that relies on spoken language processing. Speech variability across talkers provides a treasure trove of information for cognitive neuroscientists, leading to important insights into the cognitive mechanisms underlying language processing and potentially providing early signs of brain dysfunction. Current studies of speech are hamstrung by analyses that require preselecting specific temporal scales and acoustic dimensions. We propose a radically different approach: using unsupervised deep learning to discover a representational space for analysis of acoustic variation. To test this highly general approach, this method will be compared to current state-of-the art methods for analyzing individual variation in bilingual speech. This includes using the acoustic variation in second language speech to predict intelligibility and to detect difficulties in code-switching, particularly the challenges faced by individuals with Alzheimer’s Disease. The results will inform development of deep learning and cognitive neuroscience. The machine learning algorithm is fully general; it can be applied to speech from any language or any domain of language usage, expanding the range of populations and contexts that can be served by speech technology or studied by cognitive neuroscientists. The project’s integrative approach will allow computer scientists to advance our understanding of the extent to which modern deep learning architectures do or do not approximate human speech processing and allow cognitive neuroscientists to further our understanding of how meaningful acoustic distinctions are represented in speech perception and production. human speech representation. This project is funded by the Integrative Strategies for Understanding Neural and Cognitive Systems (NCS) program, which is jointly supported by the Directorates for Computer and Information Science and Engineering (CISE), Education and Human Resources (EHR), Engineering (ENG), and Social, Behavioral, and Economic Sciences (SBE).This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1016/j.jml.2023.104410
发表时间: 2023-01-23
期刊: JOURNAL OF MEMORY AND LANGUAGE
影响因子: 4.3
作者: [Goldrick, Matthew, Gollan, Tamar H.]
通讯作者: Gollan, Tamar H.
Advancement of phonetics in the 21st century: Exemplar models of speech production
21 世纪语音学的进步:语音产生的范例模型
DOI: --
发表时间: 2023
期刊: Journal of Phonetics
影响因子: 1.9
作者: [Goldrick, Matthew, Cole, Jennifer]
通讯作者: Cole, Jennifer
Doctoral Dissertation Research: The effects of experience and attitudes on heritage bilinguals' language processing
  • 批准号:
    2141430
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.35万
  • 财政年份:
    2022
  • 负责人:
    Matthew Goldrick
  • 依托单位:
Doctoral Dissertation Research: Role of Prior Knowledge in Consolidation of Novel Phonotactic Patterns for Speech Production
  • 批准号:
    2116802
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.1万
  • 财政年份:
    2021
  • 负责人:
    Matthew Goldrick
  • 依托单位:
Doctoral Dissertation Research: Why adapt? Phonotactic learning as non-native language adaptation
  • 批准号:
    1728173
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.01万
  • 财政年份:
    2017
  • 负责人:
    Matthew Goldrick
  • 依托单位:
Doctoral Dissertation Research on the Role of Domain-General Executive Functions in Language Production: Resolving conflict in lexical selection
  • 批准号:
    1420820
  • 项目类别:
    Standard Grant
  • 资助金额:
    $0.92万
  • 财政年份:
    2014
  • 负责人:
    Matthew Goldrick
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
HIF-1α调控软骨细胞衰老在骨关节炎进展中的作用及机制研究
  • 批准号:
    82371603
  • 项目类别:
    面上项目
  • 资助金额:
    49.00万元
  • 批准年份:
    2023
  • 负责人:
    陈晓
  • 依托单位:
超声驱动压电效应激活门控离子通道促眼眶膜内成骨的作用及机制研究
  • 批准号:
    82371103
  • 项目类别:
    面上项目
  • 资助金额:
    49.00万元
  • 批准年份:
    2023
  • 负责人:
    阮静
  • 依托单位:
Lienard系统的不变代数曲线、可积性与极限环问题研究
  • 批准号:
    12301200
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30.00万元
  • 批准年份:
    2023
  • 负责人:
    钱欣洁
  • 依托单位: