CompCog: HNDS-R: Self-Supervision of Visual Learning From Spatiotemporal Context
CompCog: HNDS-R: Self-Supervision of Visual Learning From Spatiotemporal Context
批准号:
2216127
负责人:
Bradley Wyble
金额:
$49.7万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-15 至 2025-08-31
中文摘要
现代计算机视觉模型是在数十亿张图像上进行训练的,但它们的健壮性仍然远远不如小孩子的视觉系统,后者的视觉体验范围要小得多。这个项目将利用我们对婴儿在出生后的头几年如何体验世界的理解,以开发新的方法来训练人工智能程序来解码他们从相机接收的信息。与计算机相比,儿童的一个优势是他们将视觉世界体验为一次太空之旅,而不是一系列随机收集的、无关的图像。因此,儿童有一种方法可以根据儿童对每个场景的有利位置来评估两个视觉场景的相似性。研究人员将以一个小孩在房子里移动的角度为模型生成高度逼真的场景,并将其用于开发一种计算机算法,该算法可以学习如何识别物体、表面和其他视觉概念。这项工作将为改善现实世界问题的计算机视觉提供新的见解,由于其在家用机器人、辅助机器人和自动驾驶汽车等领域的应用,该领域正在快速增长。该项目将支持跨学科的研究生和博士后培训,以及通过NeuroMatch制作广泛可用的STEM教育资源,NeuroMatch是疫情期间出现的一种暑期学校,作为接触学生的一种方式,同时招致的成本最低,并保持低碳足迹。研究人员在人类儿童学习方式的启发下,开发了一种关键的视觉学习理论,有可能重塑计算机视觉和机器学习中的学习基础。这项研究假设,人类视觉学习的一个关键因素是时空邻接性,这是一个事实,即当儿童在空间中移动时,世界中的图像是按顺序体验的。该项目有两个组成部分,旨在最终开发一种基于人类学习的视觉学习新算法。首先,将使用光线跟踪创建一个数据集,以生成照片级真实感图像序列,其方式类似于儿童体验它们。然后,这些图像将与自我监督深度学习的最新创新相结合,以确定时空图像序列如何通过图像分类和其他任务作为测试来增强计算机视觉。由此产生的算法将产生对视觉模式做出反应的人工神经网络。这些反应可以与通过fMRI测量的人脑中神经网络的反应进行比较,以通过表征相似性分析来确定序列学习机制是否比最先进的计算机视觉方法更接近人类视觉学习。此外,这种分析技术可以用作探照灯,突出大脑中与新开发的人工神经网络最相似的区域;这有助于确定不同的大脑区域对视觉学习的贡献。该项目资助的学生将在心理学和计算机科学之间进行研究,该项目还将有助于STEM教育资源的开发。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Modern computer vision models are trained on sets of images numbering in the billions and yet they are still far less robust than the visual systems of small children who have a much smaller range of visual experiences. This project will use our understanding of how infants experience the world in their first years of life in order to develop new methods of training artificial intelligence programs to decode information they receive from a camera. One advantage that children have over computers is that they experience the visual world as a journey through space, rather than as a series of randomly collected, unrelated images. Children thus have a way to evaluate the similarity of two visual scenes based on the child's vantage point for each scene. The investigators will generate highly realistic scenes modeled on the perspective of a young child moving through a house, which will be used to develop a computer algorithm that learns how to recognize objects, surfaces, and other visual concepts. The work will provide new insights into improving computer vision for real-world problems, a field that is under rapid growth due to its application in areas including household robots, assistive robots, and self-driving cars. The project will support interdisciplinary graduate and postdoctoral training as well as production of widely accessible STEM educational resources through Neuromatch, which is a summer school that emerged during the pandemic as a way to reach students while incurring minimal cost and maintaining a low carbon footprint. The investigators develop a critical theory of visual learning, inspired by how human children learn, with the potential to reshape the fundamentals of learning in computer vision and machine learning. The research hypothesizes that a key ingredient in human visual learning is spatiotemporal contiguity, which is the fact that images in the world are experienced in a sequence as a child moves through space. The project has two components aimed at ultimately developing a new algorithm for visual learning based on human learning. First, a data set will be created using ray-tracing to generate sequences of photorealistic images in a similar way that a child would experience them. Then, these images will be coupled with recent innovations in self-supervised deep learning to determine how spatiotemporal image sequences can augment computer vision using image classification and other tasks as tests. The resulting algorithm will produce artificial neural networks that respond to visual patterns. Those responses can be compared with the responses of neural networks in the human brain as measured through fMRI to determine through representational-similarity analysis if the sequence-learning mechanism is a better approximation of human visual learning than state-of-the-art computer vision methods. Moreover, this analysis technique can be used as a searchlight to highlight the regions in the brain that are most similar to the newly developed artificial neural networks; this is helpful for determining how different brain areas contribute to visual learning. Students supported by this project will conduct research at the interface between psychology and computer science and the project will also contribute to the development of STEM educational resources.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CompCog: Bridging the gap between behavioral and neural correlates of attention using a computational model of neural mechanisms
-
批准号:1734220
-
项目类别:Standard Grant
-
资助金额:$39.07万
-
财政年份:2017
-
负责人:Bradley Wyble
-
依托单位:
Integrating Spatial and Temporal Models of Visual Attention
-
批准号:1331073
-
项目类别:Standard Grant
-
资助金额:$33.52万
-
财政年份:2013
-
负责人:Bradley Wyble
-
依托单位:
海外基金