DiverseNet: When One Right Answer is not Enough

DiverseNet: When One Right Answer is not Enough
复制标题

DOI:
10.1109/cvpr.2018.00587
复制
发表时间:
2018-03
期刊:
2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Michael Firman;N. Campbell;L. Agapito;G. Brostow
Michael Firman;N. Campbell;L. Agapito;G. Brostow
中科院分区:
其他
文献类型:
--
作者:
Michael Firman;N. Campbell;L. Agapito;G. Brostow

文献摘要

被引文献

相似文献

机器视觉中的许多结构化预测任务都有一系列可接受的答案,而不是一个明确的真实答案。例如,图像分割受到人类标签偏见的影响。同样,有多个可能的像素值可以合理地完成被遮挡的图像区域。最先进的监督学习方法通常被优化为对每个查询进行单个测试时间预测,而无法在输出空间中找到其他模式。允许采样的现有方法往往会牺牲速度或准确性。我们介绍了一种训练神经网络的简单方法,该方法可以对每个测试时间查询进行不同的结构化预测。对于一个单一的输入,我们学会预测一系列可能的答案。与通过网络集合寻求多样性的方法相比,我们更有优势。这种随机选择学习面临模式崩溃,其中一个或多个集合成员无法接收任何训练信号。我们最好的解决方案可以部署到各种任务中,并且只涉及对现有单模体系结构、损失函数和训练机制的小修改。我们证明了我们的方法可以在三个具有挑战性的任务中进行定量改进:2D图像补全,3D体积估计和流量预测。
Many structured prediction tasks in machine vision have a collection of acceptable answers, instead of one definitive ground truth answer. Segmentation of images, for example, is subject to human labeling bias. Similarly, there are multiple possible pixel values that could plausibly complete occluded image regions. State-of-the art supervised learning methods are typically optimized to make a single test-time prediction for each query, failing to find other modes in the output space. Existing methods that allow for sampling often sacrifice speed or accuracy. We introduce a simple method for training a neural network, which enables diverse structured predictions to be made for each test-time query. For a single input, we learn to predict a range of possible answers. We compare favorably to methods that seek diversity through an ensemble of networks. Such stochastic multiple choice learning faces mode collapse, where one or more ensemble members fail to receive any training signal. Our best performing solution can be deployed for various tasks, and just involves small modifications to the existing single-mode architecture, loss function, and training regime. We demonstrate that our method results in quantitative improvements across three challenging tasks: 2D image completion, 3D volume estimation, and flow prediction.