课题基金 / 基金详情

CAREER: Visual Learning in an Open and Continual World

CAREER: Visual Learning in an Open and Continual World
职业:开放和持续世界中的视觉学习
批准号:
2239292
负责人:
Zsolt Kira
金额:
$53.51万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-02-01 至 2028-01-31

项目摘要

项目成果

Zsolt Kira的其他基金

相似基金

相关文献

中文摘要
翻译
计算机视觉领域在过去十年中取得了重大进展:这些模型现在能够有效地处理复杂图像并自动提取信息,例如检测图像中存在的对象类型以及它们的位置。然而,当前的方法需要图像中的对象类别的预先指定的列表。当系统部署在现实环境中时,这种要求是不现实的,例如自动驾驶汽车或大型照片集。如果出现新类型的对象,当前的系统将需要人类识别新对象并注释图像,然后通过一个需要大量计算资源的过程重新训练计算机视觉模型。与人类不同,系统无法自动理解图像中何时出现新类型的对象,它们与系统已经知道的对象之间的关系,以及如何在几乎没有人类注释的情况下不断更新其知识。因此,该项目旨在使计算机视觉系统能够连续自动地检测和发现新类别,并更新其模型,几乎没有人工注释。这种能力将在一系列应用中产生影响,包括照片集的个性化分析、家用机器人、自动驾驶汽车和医学成像,其中新的未知物体通常会导致误导或不正确的物体检测。该项目将通过一系列研究创新以及一些推广活动来解决这个问题,包括通过与K-12教育工作者合作来教授我们的开源课程材料,从而使人工智能教育民主化。为此,该项目的目标是为开放世界和持续学习系统创建一个框架,该框架开发了自然理解和处理不同类型分布变化的原则方法,以及在未标记数据中出现时逐渐发现和学习新类别,并将其置于丰富的语义层次结构中。这将通过首先检测可能发生的不同类型的分布偏移(例如,由于天气或全新物体的存在而引起的外观变化),并开发原则性的分布外检测和校准方法来解开它们。这些方法将用于了解它们如何影响模型的预测。随后,这种对分布变化的细粒度理解将支持增量地更新模型作为响应,而不仅仅是检测是否存在新类别并抛出结果数据。这将通过开发方法来建立长期的表示和分类器,发现新的类别,并将它们放置在一个丰富的层次语义结构。最后,半监督式持续学习将被用来逐步改进表示,并自动学习分类和检测模型,使用不同时间出现的标记和未标记数据的混合,同时避免灾难性遗忘。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The field of computer vision has seen significant progress in the past decade: These models are now able to efficiently process complex images and automatically extract information, such as detecting what type of objects exist in the image and where they are located. However, current methods require a pre-specified list of object categories that are in the images. This requirement that is unrealistic when systems are deployed in real-world contexts, such as on self-driving cars or large photo collections. If new types of objects appear, current systems will need to have humans identify the new objects and annotate the images and then retrain the computer vision model through a process that takes significant computational resources. Unlike humans, the system cannot automatically understand when new types of objects are in the images, how they relate to objects that the system already knows about, and how to continually update its knowledge given little to no human annotation. This project therefore seeks to enable a computer vision system that can continuously and automatically detect and discover new categories, as well as update its model, with little to no human annotation. Such a capability would have implications in a range of applications, including personalized analysis of photo collections, home robotics, self-driving cars, and medical imaging, where novel unknown objects often lead to misleading or incorrect object detection. The project will address this through a range of research innovations as well as through several outreach activities, including democratizing AI education by working with educators from K-12 and up to teach our open-source course materials. Towards this end, the goal of this project is to create a framework for an open-world and continual learning system that develops principled methods for naturally understanding and handling different types of distribution shifts, as well as incrementally discovering and learning new categories as they appear in unlabeled data, and placing them within a rich semantic hierarchical structure. This will be accomplished by first detecting different types of distribution shift that can occur (e.g., changes in appearance due to weather or existence of entirely new objects) and developing principled out-of-distribution detection and calibration methods to disentangle them. These methods will be used to understand how they affect the model's predictions. Subsequently, rather than just detecting whether new categories exist and throwing the resulting data out, this fine-grained understanding of distribution shift will support incrementally updating the model in response. This will be done by developing methods to build long-term representations and classifiers that discover new categories and place them within a rich hierarchical semantic structure. Finally, semi-supervised continual learning will be leveraged to incrementally refine the representations and automatically learn classification and detection models, using a mixture of labeled and unlabeled data appearing at different times, while avoiding catastrophic forgetting.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.48550/arxiv.2306.09970
发表时间: 2023-06
期刊: ArXiv
影响因子: --
作者: [Shaunak Halbe;James Smith;Junjiao Tian;Z. Kira]
通讯作者: Shaunak Halbe;James Smith;Junjiao Tian;Z. Kira
DOI: 10.1109/cvpr52729.2023.01146
发表时间: 2022-11
期刊: 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者: [James Smith;Leonid Karlinsky;V. Gutta;Paola Cascante-Bonilla;Donghyun Kim-;Assaf Arbelle;Rameswar Panda;R. Feris;Z. Kira]
通讯作者: James Smith;Leonid Karlinsky;V. Gutta;Paola Cascante-Bonilla;Donghyun Kim-;Assaf Arbelle;Rameswar Panda;R. Feris;Z. Kira
NRI: Large-Scale Collaborative Semantic Mapping using 3D Structure from Motion
国内基金
海外基金
基于多幅图象的Visual Hull重构及表面属性建模算法研究
  • 批准号:
    60373031
  • 项目类别:
    面上项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2003
  • 负责人:
    陈越
  • 依托单位: