课题基金 / 基金详情

Collaborative Research: RI: Medium: Lie group representation learning for vision

Collaborative Research: RI: Medium: Lie group representation learning for vision
协作研究:RI:中:视觉的李群表示学习
批准号:
2313150
负责人:
Nina Miolane
金额:
$50.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2026-09-30

项目摘要

项目成果

Nina Miolane的其他基金

相似基金

相关文献

中文摘要
翻译
建造能够感知、理解并在其环境中行动的智能机器的探索是我们这个时代最大的科学挑战之一。尽管人工智能(AI)取得了最新进展,但实现强大的自主视觉系统来理解物理世界并与之交互仍然是一个难题。从数学上讲,视觉需要理解各种各样的物体形状之间的关系,每个形状都受到各种各样的几何和照明变换的影响,从而导致可能的视觉场景的爆炸。该项目旨在通过开发基于数学的视觉计算理论来突破这一障碍,这将使一类新的神经网络学习算法能够将视觉场景解析为其组成对象和变换,从而使计算机能够更好地表示周围的世界。从这项研究中产生的结果和计算工具将通过课程、研讨会、黑客马拉松和贡献给Geomstats图书馆的开源软件传播给科学界和公众。这个项目的前提是,人工智能和计算机视觉目前的局限性可以用一个适当的数学框架来解决,Lie理论,它模拟了视觉世界中自然变换的层次结构。 研究人员将通过编码在可学习的G-模块(群模块)中的显式李群运算来开发基础信号处理变换的泛化。 这些模块通过将图像分解为形状及其底层转换来直接解决视觉中的组合爆炸。 具体来说,该团队将开发G模块,学习自然图像中包含的变换的群等变表示(目标1),通过仅针对特定变换折叠群轨道来鲁棒地表示形状(目标2),以及通过因子分解来解开变换和形状(目标3)。这些模块被组装成分层架构,可以学习转换和形状的复杂表示(目标4)。这些目标共同提供了一种新的范式,为现有的视觉模型奠定了基础,并为未来深度学习架构的设计提供了一套指导原则,增强了感知和理解世界的能力。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响力评审标准进行评估,被认为值得支持。
英文摘要
The quest to build intelligent machines capable of sensing, understanding and acting in their environment presents one of the great scientific challenges of our time. Despite recent advances in artificial intelligence (AI), the realization of robust, autonomous vision systems that understand and interact with the physical world remains elusive. Mathematically, vision requires understanding the relationships among an immense variety of object shapes, each subject to an immense variety of geometric and lighting transformations, leading to an explosion of possible visual scenes. This project aims to break through this barrier by developing a mathematically grounded computational theory of vision that will enable a new class of neural network learning algorithms to parse visual scenes into their constituent objects and transformations, thereby enabling computers to better represent the world around them. The results and computational tools arising from this research will be disseminated to the scientific community and general public through courses, seminars, hackathons, and open-source software contributed to the Geomstats library.The premise of this project is that the current limitations of AI and computer vision can be addressed with an appropriate mathematical framework, Lie theory, that models the hierarchical structure of natural transformations in the visual world. The investigators will develop generalizations of foundational signal processing transforms through explicit Lie group operations encoded in learnable G-Modules (Group-Modules). These modules directly tackle the combinatoric explosion in vision by factorizing images into shapes and their underlying transformations. Specifically, the team will develop G-modules that learn group-equivariant representations of the transformations contained in natural images (Aim 1), robust representations of shape by collapsing group orbits only with respect to specific transformations (Aim 2), and disentangling of transformation and shape via factorization (Aim 3). The modules are assembled into hierarchical architectures that can learn complex representations of transformations and shapes (Aim 4). Together, these aims provide a new paradigm that grounds existing models of vision and gives a set of guiding principles for the design of future deep learning architectures with enhanced abilities to sense and understand the world.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Advancing Shape Learning for Biosciences
Collaborative Research: A Unifying Deep Learning Framework Using Cell Complex Neural Networks
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)