Task2Sim: Towards Effective Pre-training and Transfer from Synthetic Data

Task2Sim: Towards Effective Pre-training and Transfer from Synthetic Data
复制标题

DOI:
10.1109/cvpr52688.2022.00898
复制
发表时间:
2021-11
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Samarth Mishra;Rameswar Panda;Cheng Perng Phoo;Chun-Fu Chen;Leonid Karlinsky;Kate Saenko;Venkatesh Saligrama;R. Feris
Samarth Mishra;Rameswar Panda;Cheng Perng Phoo;Chun-Fu Chen;Leonid Karlinsky;Kate Saenko;Venkatesh Saligrama;R. Feris
中科院分区:
其他
文献类型:
--
作者:
Samarth Mishra;Rameswar Panda;Cheng Perng Phoo;Chun-Fu Chen;Leonid Karlinsky;Kate Saenko;Venkatesh Saligrama;R. Feris

文献摘要

相似文献

Imagenet或其他真实的图像的大规模数据集上的预训练模型已经导致了计算机视觉的重大进步,尽管伴随着与管理成本,隐私,使用权和道德问题相关的缺点。在本文中,我们第一次研究了基于图形模拟器生成的合成数据的预训练模型到来自不同领域的下游任务的可移植性。在使用这种合成数据进行预训练时,我们发现不同任务的下游性能受到模拟参数(例如照明,对象姿势,背景等)的不同配置的影响,没有放之四海而皆准的解决办法。因此,最好将合成的预训练数据定制为特定的下游任务,以获得最佳性能。我们引入了Task2Sim,这是一个统一的模型,它将下游任务表示映射到最佳仿真参数,以生成合成的预训练数据。Task2Sim通过训练来学习此映射,以找到一组“可见”任务的最佳参数集。一旦经过训练,它就可以用于一次性预测新的“看不见的”任务的最佳模拟参数,而无需额外的训练。给定每个类的图像数量预算,我们对20个不同的下游任务进行了广泛的实验,结果表明,Task2Sim的任务自适应预训练数据比非自适应地选择可见和不可见任务的模拟参数具有更好的下游性能。它甚至可以与Imagenet的真实的图像进行预训练相媲美。
Pre-training models on Imagenet or other massive datasets of real images has led to major advances in Computer vision, albeit accompanied with shortcomings related to curation cost, privacy, usage rights, and ethical issues. In this paper, for the first time, we study the transferability of pre-trained models based on synthetic data generated by graphics simulators to downstream tasks from very different domains. In using such synthetic data for pre-training, we find that downstream performance on different tasks are fa-vored by different configurations of simulation parameters (e.g. lighting, object pose, backgrounds, etc.), and that there is no one-size-fits-all solution. It is thus better to tailor syn-thetic pre-training data to a specific downstream task, for best performance. We introduce Task2Sim, a unified model mapping downstream task representations to optimal sim-ulation parameters to generate synthetic pre-training data for them. Task2Sim learns this mapping by training to find the set of best parameters on a set of “seen” tasks. Once trained, it can then be used to predict best simulation pa-rameters for novel “unseen” tasks in one shot, without re-quiring additional training. Given a budget in number of images per class, our extensive experiments with 20 di-verse downstream tasks show Task2Sim's task-adaptive pre-training data results in significantly better downstream per-formance than non-adaptively choosing simulation param-eters on both seen and unseen tasks. It is even competitive with pre-training on real images from Imagenet.