CompCog: Deep causal inference grounds the perception of cognitive objects in speech
CompCog: Deep causal inference grounds the perception of cognitive objects in speech
批准号:
2240349
负责人:
Khalil Iskarous
金额:
$60.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-15 至 2026-07-31
中文摘要
人工智能系统在社会中变得越来越重要,近年来其性能有了很大的提高。然而,我们仍然不了解这些系统实际上是如何工作的,以及它们如何在一定程度上模仿人类的表现。在这项工作中,通过将这些系统的内部计算与人类执行的相应计算进行比较,开发了新的方法来探测这些系统的内部工作。我们研究的具体技能是语音识别——这是一个高度复杂的过程,因为语音是一种变化丰富、信息密集、传播迅速的交流媒介。在语音识别过程中,人类大脑处理这种复杂性的方法之一是不仅参与负责听力的大脑区域,还参与对语音产生至关重要的区域。这表明人类的认知意识到世界上引起语言的系统——嘴唇、舌头和其他发声器官的运动。当前的人工智能系统是否也能对语音产生如此深刻的因果理解?在这项工作中,我们通过深入研究这些系统中人类和机器知识的数学模型来回答这个问题。我们的工作有两个主要目标。首先是技术方面:了解人工智能系统的内部工作原理,这最终是将其能力用于社会效益的必要步骤。第二个是科学方面的:在广泛应用于社会之前,认知科学家开发了人工智能系统来理解人类的认知,通过探索最先进的机器学习的内部工作原理,作为一种认知模型,我们可能能够更好地理解人类是如何感知语言的。因此,在这项工作中,科学和技术相互促进,正如它们在过去所做的那样成功。这个研究项目专门探讨了人类和计算机的语言产生和感知之间的关系。为了做到这一点,使用实时磁共振成像(MRI)对三种语言(英语,俄语,韩语)的使用者进行成像,该成像可以生动地详细显示语音发音器的运动情况。同时记录说话人的语音音频信号。使用语音产生的数学模型、现代语音识别系统和人类神经节律如何分析语音的数学模型对数据进行分析。实验操作揭示了每个系统中的表征如何与其他系统中的表征相对应。这一策略告诉我们关于人类认知的科学,同时也照亮了机器模拟人类能力的黑箱技术。在未来,除了推进科学和技术,我们预计这些知识的应用,以创造新的小型语音识别系统,可以协助濒危语言的文件。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Artificial Intelligence systems are becoming more and more important in society, and their performance has improved enormously in recent years. Yet, we still do not understand how these systems actually work, and how they emulate human performance, to the extent that they do. In this work, novel methods are developed for probing the inner working of these systems by comparing their internal computations with corresponding computations humans perform. The specific skill we probe is speech recognition—a highly complex process, as speech is a richly variable, information dense, and quickly transmitted medium of communication. One of the ways that the human brain deals with this complexity during speech recognition is by engaging not only brain areas responsible for listening but also areas crucial to the production of speech. This suggests that human cognition is aware of the systems in the world that cause speech—the movements of the lips, tongue, and other vocal organs. Do current artificial intelligence systems also develop such a deep causal understanding of speech? In this work we answer this question by delving into the mathematical models of both human and machine knowledge in these systems. Our work has two major goals. The first is technological: understanding how artificial intelligence systems actually work on the inside, which is ultimately a necessary step in directing their abilities to societal benefit. The second is scientific: before widespread use in society, artificial intelligence systems were developed by cognitive scientists to understand human cognition, and by probing the inner workings of state-of-the-art machine learning as a cognitive model we may be able to better understand how humans perceive speech. In this work, therefore, science and technology further each other, as they have done successfully in the past.This research program specifically probes the relationship between the production and perception of speech in humans and computers. To do so, speakers of three languages (English, Russian, Korean) are imaged using a real time Magnetic Resonance Imaging (MRI), which shows in vivid detail how the speech articulators move. Speaker’s speech audio signals are recorded simultaneously. The data are analyzed using mathematical models of speech production, modern speech recognition systems, and mathematical models of how human neural rhythms analyze speech. Experimental manipulations unveil how the representations in each of the systems corresponds to those in the others. This strategy inform us about the science of human cognition hand-in-hand with illuminating the black-box technology of machine emulation of the human capability. In the future, in addition to advancing science and technology, we anticipate the application of this knowledge to the creation of novel small-sized speech recognition systems that can assist in the documentation of endangered languages.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
REG: Morphological Investigations in Formosan Languages
-
批准号:1219533
-
项目类别:Standard Grant
-
资助金额:$1.26万
-
财政年份:2012
-
负责人:Khalil Iskarous
-
依托单位:
INSPIRE: Dynamical Principles of Animal Movement
-
批准号:1246750
-
项目类别:Standard Grant
-
资助金额:$97.4万
-
财政年份:2012
-
负责人:Khalil Iskarous
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
-
批准号:2026JJ81909
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:胡曦
-
依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
-
批准号:12271434
-
项目类别:面上项目
-
资助金额:46万元
-
批准年份:2022
-
负责人:贺小伟
-
依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
-
批准号:2020A151501709
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2020
-
负责人:谢怡
-
依托单位:
面向Deep Web的数据整合关键技术研究
-
批准号:61872168
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2018
-
负责人:董永权
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于语义计算的海量Deep Web知识探索机制研究
-
批准号:61272411
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:赵峰
-
依托单位:
Deep Web数据集成查询结果抽取与整合关键技术研究
-
批准号:61100167
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:董永权
-
依托单位:
面向Deep Web的大规模知识库自动构建方法研究
-
批准号:61170020
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2011
-
负责人:崔志明
-
依托单位:
Deep Web敏感聚合信息保护方法研究
-
批准号:61003054
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2010
-
负责人:赵朋朋
-
依托单位: