3DViewGraph: Learning Global Features for 3D Shapes from A Graph of Unordered Views with Attention

3DViewGraph: Learning Global Features for 3D Shapes from A Graph of Unordered Views with Attention
复制标题

DOI:
10.24963/ijcai.2019/107
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Zhizhong Han;Xiyang Wang;C. Vong;Yu-Shen Liu;Matthias Zwicker;C. L. P. Chen
Zhizhong Han;Xiyang Wang;C. Vong;Yu-Shen Liu;Matthias Zwicker;C. L. P. Chen
中科院分区:
其他
文献类型:
--
作者:
Zhizhong Han;Xiyang Wang;C. Vong;Yu-Shen Liu;Matthias Zwicker;C. L. P. Chen

文献摘要

相似文献

通过聚合多个视图的信息来学习全局特征已被证明对于 3D 形状分析是有效的。对于深度学习模型中的视图聚合,池化已被广泛应用。然而,池化会导致视图内内容以及视图之间的空间关系的丢失,从而限制了学习特征的可区分性。我们提出 3DViewGraph 来解决这个问题,它通过更有效地聚合无序视图和注意力来学习 3D 全局特征。具体来说,围绕形状采取的无序视图被视为视图图上的视图节点。 3DViewGraph 首先学习一种新颖的潜在语义映射,将低级视图特征投影到较低维空间中有意义的潜在语义嵌入,该空间由潜在语义模式跨越。然后,每对视图节点的内容和空间信息通过新颖的空间模式相关性进行编码,其中计算潜在语义模式之间的相关性。最后,所有空间模式相关性都与通过新颖的注意力机制学习到的注意力权重相结合。通过突出显示具有独特特征的无序视图节点并抑制具有外观模糊性的视图节点,进一步提高了学习特征的可辨别性。我们证明 3DViewGraph 在三个大型基准测试中优于最先进的方法。
Learning global features by aggregating information over multiple views has been shown to be effective for 3D shape analysis. For view aggregation in deep learning models, pooling has been applied extensively. However, pooling leads to a loss of the content within views, and the spatial relationship among views, which limits the discriminability of learned features. We propose 3DViewGraph to resolve this issue, which learns 3D global features by more effectively aggregating unordered views with attention. Specifically, unordered views taken around a shape are regarded as view nodes on a view graph. 3DViewGraph first learns a novel latent semantic mapping to project low-level view features into meaningful latent semantic embeddings in a lower dimensional space, which is spanned by latent semantic patterns. Then, the content and spatial information of each pair of view nodes are encoded by a novel spatial pattern correlation, where the correlation is computed among latent semantic patterns. Finally, all spatial pattern correlations are integrated with attention weights learned by a novel attention mechanism. This further increases the discriminability of learned features by highlighting the unordered view nodes with distinctive characteristics and depressing the ones with appearance ambiguity. We show that 3DViewGraph outperforms state-of-the-art methods under three large-scale benchmarks.