DANCE: DAta-Network Co-optimization for Efficient Segmentation Model Training and Inference

DANCE: DAta-Network Co-optimization for Efficient Segmentation Model Training and Inference
复制标题

DOI:
10.1145/3510835
复制
发表时间:
2021-07
期刊:
ACM Transactions on Design Automation of Electronic Systems (TODAES)
影响因子:
--
通讯作者:
Chaojian Li;Wuyang Chen;Yuchen Gu;Tianlong Chen;Yonggan Fu;Zhangyang Wang;Yingyan Lin
Chaojian Li;Wuyang Chen;Yuchen Gu;Tianlong Chen;Yonggan Fu;Zhangyang Wang;Yingyan Lin
中科院分区:
其他
文献类型:
--
作者:
Chaojian Li;Wuyang Chen;Yuchen Gu;Tianlong Chen;Yonggan Fu;Zhangyang Wang;Yingyan Lin

文献摘要

相似文献

目前,场景理解的语义分割被广泛应用,这对算法的效率提出了很大的挑战,特别是在资源有限的平台上的应用。目前的分割模型是在海量高分辨率场景图像上训练和评估的(“数据级”),并且由于所需的多尺度聚集(“网络级”)而遭受着昂贵的计算。在这两种情况下,由于分割模型通常需要大的输入分辨率和沉重的计算负担,在训练和推理中的计算和能量成本是显著的。为此,我们提出了DIDAY、通用自动数据网络联合优化算法,用于高效的分割模型训练和推理。与仅专注于轻量级网络设计的现有高效分段方法不同,DIXY通过输入数据操作和网络结构瘦身,将自己区分为自动同时进行数据网络协同优化。具体地说,Dance集成了自动数据瘦身,它自适应地对输入图像进行下采样/丢弃,并根据图像的空间复杂性控制它们对训练损失的相应贡献。这种下采样操作,除了直接缩小与输入大小相关的成本外,还缩小了输入对象和上下文尺度的动态范围,因此激励我们也自适应地细化网络以匹配下采样数据。广泛的实验和烧蚀研究(在四个SOTA分割模型和两个训练设置下的三个流行的分割数据集上)表明,DIZE可以在高效分割(降低训练成本、较低的推理代价和更好的Mean-Over-Over-Union(MIUU))方面实现“双赢”。具体而言,DIZE训练可降低↓25%-↓77%,推理↓31%-↓56%,同时提高MIU0.71%-↓13.34%。
Semantic segmentation for scene understanding is nowadays widely demanded, raising significant challenges for the algorithm efficiency, especially its applications on resource-limited platforms. Current segmentation models are trained and evaluated on massive high-resolution scene images (“data-level”) and suffer from the expensive computation arising from the required multi-scale aggregation (“network level”). In both folds, the computational and energy costs in training and inference are notable due to the often desired large input resolutions and heavy computational burden of segmentation models. To this end, we propose DANCE, general automated DAta-Network Co-optimization for Efficient segmentation model training and inference. Distinct from existing efficient segmentation approaches that focus merely on light-weight network design, DANCE distinguishes itself as an automated simultaneous data-network co-optimization via both input data manipulation and network architecture slimming. Specifically, DANCE integrates automated data slimming which adaptively downsamples/drops input images and controls their corresponding contribution to the training loss guided by the images’ spatial complexity. Such a downsampling operation, in addition to slimming down the cost associated with the input size directly, also shrinks the dynamic range of input object and context scales, therefore motivating us to also adaptively slim the network to match the downsampled data. Extensive experiments and ablating studies (on four SOTA segmentation models with three popular segmentation datasets under two training settings) demonstrate that DANCE can achieve “all-win” towards efficient segmentation (reduced training cost, less expensive inference, and better mean Intersection-over-Union (mIoU)). Specifically, DANCE can reduce ↓25%–↓77% energy consumption in training, ↓31%–↓56% in inference, while boosting the mIoU by ↓0.71%–↑ 13.34%.