Learning of Visual Relations: The Devil is in the Tails

Learning of Visual Relations: The Devil is in the Tails
复制标题

DOI:
10.1109/iccv48922.2021.01512
复制
发表时间:
2021-08
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Alakh Desai;Tz-Ying Wu;Subarna Tripathi;N. Vasconcelos
Alakh Desai;Tz-Ying Wu;Subarna Tripathi;N. Vasconcelos
中科院分区:
其他
文献类型:
--
作者:
Alakh Desai;Tz-Ying Wu;Subarna Tripathi;N. Vasconcelos

文献摘要

相似文献

最近,大量努力致力于建模视觉关系。这主要通过添加参数和增加模型复杂性来解决架构的设计。但是,由于关节推理的对象组的联合性质,视觉关系学习是一个长尾的问题。通常,由于其过度贴身的趋势,增加模型的复杂性通常不适合长尾巴问题。在本文中,我们探讨了另一种假设,表示魔鬼在尾部。在这个假设下,通过保持模型简单,但提高其应对长尾分布的能力来实现更好的性能。为了检验这一假设,我们设计了一种新的培训视觉关系模型的方法,该方法受到最先进的长尾识别文献的启发。这是基于迭代脱钩的训练计划,表示尾巴中的魔鬼的脱钩训练(DT2)。 DT2采用了一种新型的采样方法,即交替的类平衡采样(ACB)来捕获长尾实体和视觉关系的谓词分布之间的相互作用。结果表明,借助非常简单的体系结构,DT2-ACB的表现明显超出了场景图生成任务上更复杂的最新方法。这表明,必须考虑与问题的长尾巴性质一起考虑复杂模型的发展。
Significant effort has been recently devoted to modeling visual relations. This has mostly addressed the design of architectures, typically by adding parameters and increasing model complexity. However, visual relation learning is a long-tailed problem, due to the combinatorial nature of joint reasoning about groups of objects. Increasing model complexity is, in general, illsuited for long-tailed problems due to their tendency to overfit. In this paper, we explore an alternative hypothesis, denoted the Devil is in the Tails. Under this hypothesis, better performance is achieved by keeping the model simple but improving its ability to cope with long-tailed distributions. To test this hypothesis, we devise a new approach for training visual relationships models, which is inspired by state-of-the-art long-tailed recognition literature. This is based on an iterative decoupled training scheme, denoted Decoupled Training for Devil in the Tails (DT2). DT2 employs a novel sampling approach, Alternating Class-Balanced Sampling (ACBS), to capture the interplay between the long-tailed entity and predicate distributions of visual relations. Results show that, with an extremely simple architecture, DT2-ACBS significantly out-performs much more complex state-of-the-art methods on scene graph generation tasks. This suggests that the development of sophisticated models must be considered in tandem with the long-tailed nature of the problem.