Latent transformations neural network for object view synthesis

Latent transformations neural network for object view synthesis
复制标题

DOI:
10.1007/s00371-019-01755-x
复制
发表时间:
2019-10
期刊:
The Visual Computer
影响因子:
--
通讯作者:
Sangpil Kim;Nick Winovich;Hyung-gun Chi;Guang Lin;K. Ramani
Sangpil Kim;Nick Winovich;Hyung-gun Chi;Guang Lin;K. Ramani
中科院分区:
其他
文献类型:
--
作者:
Sangpil Kim;Nick Winovich;Hyung-gun Chi;Guang Lin;K. Ramani

文献摘要

相似文献

我们提出了一个完全卷积的条件生成神经网络,潜在的转换神经网络,能够刚性和非刚性对象视图合成使用一个轻量级的架构,适合于实时应用和嵌入式系统。在现有的对象视图合成方法,通过级联,将条件信息,我们引入了一个专用的网络组件,条件变换单元。该单元被设计为学习对应于指定目标视图的潜在空间变换。此外,定义了一致性损失项以引导网络学习所需的潜在空间映射,构建了任务划分的解码器以改进生成的对象视图的质量,并引入了自适应解码器以改善对抗训练过程。所提出的方法的通用性证明了三个不同的任务的集合:多视图合成真实的手深度图像,视图合成的真实的和合成的脸,和刚性物体的旋转。所提出的模型被证明是可比的最先进的方法在结构相似性指数的措施和度量,同时实现了减少24%的计算时间推断的新的图像。
We propose a fully convolutional conditional generative neural network, the latent transformation neural network, capable of rigid and non-rigid object view synthesis using a lightweight architecture suited for real-time applications and embedded systems. In contrast to existing object view synthesis methods which incorporate conditioning information via concatenation, we introduce a dedicated network component, the conditional transformation unit. This unit is designed to learn the latent space transformations corresponding to specified target views. In addition, a consistency loss term is defined to guide the network toward learning the desired latent space mappings, a task-divided decoder is constructed to refine the quality of generated views of objects, and an adaptive discriminator is introduced to improve the adversarial training process. The generalizability of the proposed methodology is demonstrated on a collection of three diverse tasks: multi-view synthesis on real hand depth images, view synthesis of real and synthetic faces, and the rotation of rigid objects. The proposed model is shown to be comparable with the state-of-the-art methods in structural similarity index measure andmetrics while simultaneously achieving a 24% reduction in the compute time for inference of novel images.