CAREER: Differentiable Programming for Visual Computing
CAREER: Differentiable Programming for Visual Computing
批准号:
2238839
负责人:
Tzu-Mao Li
金额:
$61.51万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-03-01 至 2028-02-29
中文摘要
深度神经网络是一种现代机器学习方法,在处理图像和3D内容等视觉数据方面可以产生出色的结果。然而,深度神经网络有一些已知的局限性。它们通常不直接建模底层物理过程(例如,光传输或动力学),它们需要大量的计算资源进行训练和推理,并且难以调试和控制。另一方面,虽然经典的视觉计算算法明确地模拟视觉数据的形成(例如,相机如何捕捉图像或物体如何物理移动)受到这些问题的影响较小,但它们通常不像现代机器学习方法那样广泛应用,因为它们没有从大量的经验中学习。本研究将通过创建可微的经典视觉计算算法来弥合这两种方法之间的差距。也就是说,这些算法的功能平滑地依赖于一组内部参数,这些参数可以使用深度学习方法自动调整。该项目将使用数据优化这些特定领域的可微分视觉计算程序,以获得两全面性的优势。项目成果将对应用产生广泛影响,例如使自动驾驶汽车做出更好的决策,训练机器人使用物理信息与环境交互,创建更逼真的虚拟世界,设计具有更好照明的建筑物,设计具有理想外观和功能的物理对象,以及允许电影艺术家创作更好的电影镜头。通过这项研究开发的系统将被纳入新的编程课程和教程中,PI致力于与早期职业学者项目合作,以促进视觉计算和可微分编程的参与。这个项目追求一个协同计划,包括可微分编程系统、算法和应用程序的设计。为此,有必要采用特定领域的算法,并以正确有效的方式计算其导数,设计新的视觉计算算法,利用数据驱动的先验和特定领域的知识,并将问题参数化以进行优化,以避免局部最小值并满足约束。传统的自动微分和现代深度学习系统都无法解决这些挑战。算法、系统和应用程序将共同发展,互相帮助。具体而言,该项目将开发可微分编程语言,可以适当地处理不连续,并通过利用可视化计算程序中的结构化稀疏性自动优化代码性能,以有效地处理数百万或数十亿像素,粒子或三角形。它还将开发新的领域特定的可微分视觉计算算法,通过保留经典算法的结构,同时用数据驱动的元素取代算法的手工构建的启发式组件,提高图像处理和物理模拟的效率和准确性。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Deep neural networks are a modern machine learning method that is known to produce excellent results for processing visual data such as images and 3D content. However, deep neural networks have some known limitations. They usually do not model the underlying physical process (e.g., light transport or dynamics) directly, they require significant computational resources for training and inference, and they are difficult to debug and control. On the other hand, while classical visual computing algorithms that explicitly model the formation of visual data (e.g., how a camera captures a picture or how objects move physically) suffer less from these issues, they often do not apply as broadly as modern machine learning methods because they do not learn from a large amount of experience. This research will bridge the gap between the two approaches by creating classical visual computing algorithms that are differentiable. That is, the functioning of these algorithms depends smoothly on a set of internal parameters that can be tuned automatically using deep learning approaches. The project will optimize these domain-specific differentiable visual computing programs using data to get the best of both worlds. Project outcomes will have broad impact in applications such as enabling self-driving cars to make better decisions, training robots to interact with the environment using physical information, creating more realistic virtual worlds, designing buildings with better lighting, designing physical objects with desired appearance and functionality, and allowing movie artists to create better film shots. The systems developed through this research will be incorporated into new programming courses and tutorials, and the PI is committed to working with early career scholar programs to promote participation in visual computing and differentiable programming.This project pursues a synergistic plan that includes the design of differentiable programming systems, algorithms, and applications. To these ends, it will be necessary to adapt domain-specific algorithms and compute their derivatives in correct and efficient manners, to design new visual computing algorithms that leverage both data-driven priors and domain-specific knowledge, and to parameterize the problem for optimization to avoid local minima and satisfy constraints. Neither traditional automatic differentiation nor modern deep learning systems address these challenges. The algorithms, systems, and applications will evolve together to help each other. Concretely, this project will develop differentiable programming languages that can properly handle discontinuities, and automatically optimize code performance to efficiently process millions or billions of pixels, particles, or triangles by exploiting structured sparsity in visual computing programs. It will also develop new domain-specific differentiable visual computing algorithms with improved efficiency and accuracy in image processing and physical simulation, by retaining the structures of classical algorithms while replacing hand-built heuristic components of the algorithm with data-driven elements.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3618379
发表时间:
2023-12
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
作者:
[Yash Belhe;Michaël Gharbi;Matthew Fisher;Iliyan Georgiev;Ravi Ramamoorthi;Tzu-Mao Li]
通讯作者:
Yash Belhe;Michaël Gharbi;Matthew Fisher;Iliyan Georgiev;Ravi Ramamoorthi;Tzu-Mao Li
DOI:
10.1145/3606938
发表时间:
2023-08
期刊:
Proceedings of the ACM on Computer Graphics and Interactive Techniques
影响因子:
1.3
作者:
[Shiyang Jia;Stephanie Wang;Tzu-Mao Li;Albert Chern]
通讯作者:
Shiyang Jia;Stephanie Wang;Tzu-Mao Li;Albert Chern
Inferring the Future by Imagining the Past
通过想象过去来推断未来
DOI:
--
发表时间:
2023
期刊:
Conference on Neural Information Processing Systems
影响因子:
--
作者:
[Kartik Chandra, Tony Chen]
通讯作者:
Kartik Chandra, Tony Chen
DOI:
10.1145/3618353
发表时间:
2023-12
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
作者:
[Sai Praveen Bangaru;Lifan Wu;Tzu-Mao Li;Jacob Munkberg;Gilbert Bernstein;Jonathan Ragan-Kelley;Frédo Durand;Aaron E. Lefohn;Yong He]
通讯作者:
Sai Praveen Bangaru;Lifan Wu;Tzu-Mao Li;Jacob Munkberg;Gilbert Bernstein;Jonathan Ragan-Kelley;Frédo Durand;Aaron E. Lefohn;Yong He
DOI:
10.1145/3588432.3591510
发表时间:
2023
期刊:
ACM
影响因子:
--
作者:
[Chandra, Kartik, Li, Tzu-Mao, Tenenbaum, Joshua, Ragan-Kelley, Jonathan]
通讯作者:
Ragan-Kelley, Jonathan
海外基金