Making Machine Learning on Static and Dynamic 3D Data Practical
Making Machine Learning on Static and Dynamic 3D Data Practical
批准号:
405799936
负责人:
Professor Dr.-Ing. Matthias Nießner
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2019
资助国家:
德国
项目状态:
已结题
起止时间:
2018-12-31 至 2022-12-31
中文摘要
在过去五年中,深度学习的进展导致在允许计算机通过视觉输入理解真实世界方面取得了重大进展,从而打开了从机器人到虚拟和增强现实以及医疗和工业4.0应用的许多机会。这些机器学习结构大多是卷积神经网络(CNN),它能够从图像中学习强大的特征,甚至使用生成性对抗网络(GAN)从头开始生成高度逼真的图像。在2D图像领域,我们已经看到了识别和生成任务的巨大成功。不幸的是,对于3D数据,例如自动驾驶汽车的3D扫描获得的数据,研究还处于起步阶段。这个3D方向需要进一步的探索,因为我们的世界本质上是三维的(例如人类用两只眼睛看),当考虑到时间域时,甚至是四维的。事实上,在3D中执行场景理解具有显著的优势;例如,机器学习方法不需要学习视点不变性,因此需要更少的训练数据。然而,额外的三维(对于动力学来说是第四维)带来了巨大的计算和存储开销,这一直是这些应用程序的主要瓶颈。在该建议中,我们通过开发用于3D和4D数据分析的高效机器学习算法来解决这一缺点。特别是,我们将开发深度学习体系结构和训练方法,能够有效地对不同类型的静态和动态3D数据表示进行建模,包括体素体积、RGB-D图像、点云、多视点图像和网格上的稀疏空间和时间表示。我们将进一步构建为我们的场景设计的新数据集,这些数据集从真实世界捕获,以及通过模拟渲染合成生成,并进行增强以缩小人工数据和真实数据之间的现实差距。最后,我们将开发新的神经网络体系结构,设计用于嵌入空间和特定时间域的辨别性和生成性应用程序。为了展示我们的学习方法,我们将它们应用于静态和动态的3D重建任务,以及3D和4D的语义场景理解,重点是融合空间和时间域。
英文摘要
In the last five years, advances in deep learning have led to significant progress in allowing computers to understand the real world from visual input, thus opening up many opportunities ranging from robotics to virtual and augmented reality, as well as medical and industry 4.0 applications. Most of these machine learning architectures are convolutional neural networks (CNNs), which are able to learn powerful features from images, and even generate highly-realistic pictures from scratch using generative adversarial networks (GANs). In the 2D image domain, we have seen tremendous success in both discriminative and generative tasks.Unfortunately, for 3D data, e.g. data obtained from 3D scans on autonomous cars, research is only at the infancy. This 3D direction requires further exploration, as our world is inherently three-dimensional (e.g. humans see with two eyes), and even four-dimensional when considering the temporal domain. In fact, performing scene understanding in 3D has significant advantages; for instance, a machine learning approach does not need to learn viewpoint invariance, and thus requires less training data. However, the additional third dimension (and fourth for dynamics) comes at significant computational and memory overhead, which has so far been the major bottleneck in these applications.In this proposal, we address this shortcoming by developing efficient machine learning algorithms for 3D and 4D data analysis. In particular, we will develop deep learning architectures and training methods capable of efficiently modeling different types of static and dynamic 3D data representations, including sparse sparse spatial and temporal representations on voxel volumes, RGB-D images, point clouds, multi-view images, and meshes. We will further construct new datasets designed for our scenario, captured from the real-world, as well as synthetically generated with simulated renderings, augmented to reduce the reality gap between artificial and real data. Finally, we will develop new neural network architectures designed for discriminative and generative applications embedded in spatial and specifically temporal domains. In order to showcase our learning methods, we will apply them to static and dynamic 3D reconstruction tasks, as well as semantic scene understanding in 3D and 4D with an emphasis on fusing the spatial and temporal domains.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Domain Transfer with Generative Models and Neural Rendering
-
批准号:453990920
-
项目类别:Research Units
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Professor Dr.-Ing. Matthias Nießner
-
依托单位:
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位: