CompCog: HNDS-R: Self-Supervision of Visual Learning From Spatiotemporal Context
CompCog: HNDS-R: Self-Supervision of Visual Learning From Spatiotemporal Context
批准号:
2216127
负责人:
Bradley Wyble
金额:
$49.7万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-15 至 2025-08-31
中文摘要
现代计算机视觉模型是在数以十亿计的图像集上进行训练的,但它们仍然远不如视觉经验范围小得多的小孩子的视觉系统健壮。这个项目将利用我们对婴儿在生命最初几年如何体验世界的理解,来开发训练人工智能程序的新方法,以解码他们从相机接收到的信息。与电脑相比,孩子们的一个优势是,他们体验的视觉世界是一场穿越空间的旅行,而不是一系列随机收集的、不相关的图像。这样,孩子们就有了一种方法来评估两个视觉场景的相似性,这是基于孩子对每个场景的有利位置。研究人员将以一个小孩在房子里移动的视角为模型,生成高度逼真的场景,这些场景将用于开发一种计算机算法,该算法可以学习如何识别物体、表面和其他视觉概念。这项工作将为改善现实世界问题的计算机视觉提供新的见解,这一领域由于在家用机器人、辅助机器人和自动驾驶汽车等领域的应用而迅速发展。该项目将支持跨学科的研究生和博士后培训,并通过Neuromatch提供广泛获取的STEM教育资源。Neuromatch是在疫情期间出现的一所暑期学校,旨在以最低成本和保持低碳足迹的方式接触学生。受人类儿童学习方式的启发,研究人员开发了一种视觉学习的批判性理论,有可能重塑计算机视觉和机器学习的学习基础。该研究假设,人类视觉学习的一个关键因素是时空连续性,即当儿童在空间中移动时,世界上的图像是按顺序体验的。该项目有两个组成部分,旨在最终开发一种基于人类学习的视觉学习新算法。首先,将使用光线追踪创建一个数据集,以类似于儿童体验的方式生成逼真的图像序列。然后,这些图像将与自我监督深度学习的最新创新相结合,以确定时空图像序列如何使用图像分类和其他任务作为测试来增强计算机视觉。由此产生的算法将产生对视觉模式做出反应的人工神经网络。这些反应可以与通过fMRI测量的人脑神经网络的反应进行比较,通过表征相似性分析来确定序列学习机制是否比最先进的计算机视觉方法更接近人类视觉学习。此外,这种分析技术可以作为探照灯来突出大脑中与新开发的人工神经网络最相似的区域;这有助于确定不同的大脑区域对视觉学习的贡献。本项目资助的学生将在心理学和计算机科学的界面上进行研究,并为STEM教育资源的开发做出贡献。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Modern computer vision models are trained on sets of images numbering in the billions and yet they are still far less robust than the visual systems of small children who have a much smaller range of visual experiences. This project will use our understanding of how infants experience the world in their first years of life in order to develop new methods of training artificial intelligence programs to decode information they receive from a camera. One advantage that children have over computers is that they experience the visual world as a journey through space, rather than as a series of randomly collected, unrelated images. Children thus have a way to evaluate the similarity of two visual scenes based on the child's vantage point for each scene. The investigators will generate highly realistic scenes modeled on the perspective of a young child moving through a house, which will be used to develop a computer algorithm that learns how to recognize objects, surfaces, and other visual concepts. The work will provide new insights into improving computer vision for real-world problems, a field that is under rapid growth due to its application in areas including household robots, assistive robots, and self-driving cars. The project will support interdisciplinary graduate and postdoctoral training as well as production of widely accessible STEM educational resources through Neuromatch, which is a summer school that emerged during the pandemic as a way to reach students while incurring minimal cost and maintaining a low carbon footprint. The investigators develop a critical theory of visual learning, inspired by how human children learn, with the potential to reshape the fundamentals of learning in computer vision and machine learning. The research hypothesizes that a key ingredient in human visual learning is spatiotemporal contiguity, which is the fact that images in the world are experienced in a sequence as a child moves through space. The project has two components aimed at ultimately developing a new algorithm for visual learning based on human learning. First, a data set will be created using ray-tracing to generate sequences of photorealistic images in a similar way that a child would experience them. Then, these images will be coupled with recent innovations in self-supervised deep learning to determine how spatiotemporal image sequences can augment computer vision using image classification and other tasks as tests. The resulting algorithm will produce artificial neural networks that respond to visual patterns. Those responses can be compared with the responses of neural networks in the human brain as measured through fMRI to determine through representational-similarity analysis if the sequence-learning mechanism is a better approximation of human visual learning than state-of-the-art computer vision methods. Moreover, this analysis technique can be used as a searchlight to highlight the regions in the brain that are most similar to the newly developed artificial neural networks; this is helpful for determining how different brain areas contribute to visual learning. Students supported by this project will conduct research at the interface between psychology and computer science and the project will also contribute to the development of STEM educational resources.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CompCog: Bridging the gap between behavioral and neural correlates of attention using a computational model of neural mechanisms
-
批准号:1734220
-
项目类别:Standard Grant
-
资助金额:$39.07万
-
财政年份:2017
-
负责人:Bradley Wyble
-
依托单位:
Integrating Spatial and Temporal Models of Visual Attention
-
批准号:1331073
-
项目类别:Standard Grant
-
资助金额:$33.52万
-
财政年份:2013
-
负责人:Bradley Wyble
-
依托单位:
海外基金