Efficient Controllable Multi-Task Architectures

Efficient Controllable Multi-Task Architectures
复制标题

DOI:
10.1109/iccv51070.2023.00528
复制
发表时间:
2023-08
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Abhishek Aich;S. Schulter;A. Roy-Chowdhury;Manmohan Chandraker;Yumin Suh
Abhishek Aich;S. Schulter;A. Roy-Chowdhury;Manmohan Chandraker;Yumin Suh
中科院分区:
其他
文献类型:
--
作者:
Abhishek Aich;S. Schulter;A. Roy-Chowdhury;Manmohan Chandraker;Yumin Suh

文献摘要

相似文献

我们的目标是训练一个多任务模型,这样用户就可以在部署后调整所需的计算预算和任务性能的相对重要性,而无需重新训练。这使得能够针对动态变化的用户需求优化性能,而无需为各种场景训练和保存模型而产生大量计算开销。为此,我们提出了一个多任务模型,包括一个共享的编码器和特定于任务的解码器,编码器和解码器的通道宽度是可瘦身的。我们的主要思想是通过改变特定于任务的解码器的容量来控制任务的重要性,同时通过联合调整编码器的容量来控制总的计算成本。这通过允许针对给定预算的更强编码器来提高整体准确性,增加对计算成本的控制,并基于用户的约束提供高质量的精简子架构。我们的训练策略涉及一种新的“简化不变知识蒸馏”损失,该损失强制主干表示在不同的运行时宽度配置下保持不变,以提高准确性。此外,我们提出了一个简单但有效的搜索算法,将用户的约束转换为运行时的宽度配置的共享编码器和任务解码器,采样的子架构。搜索算法的关键规则是为更高偏好的任务解码器提供更大的计算预算,同时搜索增强整体MTL性能的共享编码器配置。三个多任务基准测试(PASCALContext,NYUDv 2和CIFAR 100-MTL)与不同的骨干架构的各种实验证明了我们的方法的优势。例如,我们的方法在NYUD-v2数据集中比以前的方法显示出更高的可控性,同时产生更少的计算成本。
We aim to train a multi-task model such that users can adjust the desired compute budget and relative importance of task performances after deployment, without retraining. This enables optimizing performance for dynamically varying user needs, without heavy computational overhead to train and save models for various scenarios. To this end, we propose a multi-task model consisting of a shared encoder and task-specific decoders where both encoder and decoder channel widths are slimmable. Our key idea is to control the task importance by varying the capacities of task-specific decoders, while controlling the total computational cost by jointly adjusting the encoder capacity. This improves overall accuracy by allowing a stronger encoder for a given budget, increases control over computational cost, and delivers high-quality slimmed sub-architectures based on user’s constraints. Our training strategy involves a novel ‘Configuration-Invariant Knowledge Distillation’ loss that enforces backbone representations to be invariant under different runtime width configurations to enhance accuracy. Further, we present a simple but effective search algorithm that translates user constraints to runtime width configurations of both the shared encoder and task decoders, for sampling the sub-architectures. The key rule for the search algorithm is to provide a larger computational budget to the higher preferred task decoder, while searching a shared encoder configuration that enhances the overall MTL performance. Various experiments on three multi-task benchmarks (PASCALContext, NYUDv2, and CIFAR100-MTL) with diverse backbone architectures demonstrate the advantage of our approach. For example, our method shows a higher controllability by ∼ 33.5% in the NYUD-v2 dataset over prior methods, while incurring much less compute cost.