Scaling up and Stabilizing Differentiable Planning with Implicit Differentiation

Scaling up and Stabilizing Differentiable Planning with Implicit Differentiation
复制标题

DOI:
10.48550/arxiv.2210.13542
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Linfeng Zhao;Huazhe Xu;Lawson L. S. Wong
Linfeng Zhao;Huazhe Xu;Lawson L. S. Wong
中科院分区:
其他
文献类型:
--
作者:
Linfeng Zhao;Huazhe Xu;Lawson L. S. Wong

文献摘要

被引文献

相似文献

可区分规划保证了端到端的差异化和适应性。然而,一个问题阻碍了它扩大到更大规模的问题:他们需要通过前向迭代层来区分,以计算梯度,这耦合了前向计算和后向传播,并且需要平衡前向计划器的性能和后向传递的计算成本。为了缓解这一问题,我们建议通过Bellman不动点方程来区分价值迭代网络及其变体的正向和反向传递,从而实现恒定的反向成本(在计划范围内)和灵活的向前预算,并有助于扩展到大型任务。我们研究了隐式版本VIN及其变体的收敛稳定性、可扩展性和效率,并展示了它们在一系列规划任务中的优势:二维导航、视觉导航和在配置空间和工作空间中的两自由度操纵。
Differentiable planning promises end-to-end differentiability and adaptivity. However, an issue prevents it from scaling up to larger-scale problems: they need to differentiate through forward iteration layers to compute gradients, which couples forward computation and backpropagation, and needs to balance forward planner performance and computational cost of the backward pass. To alleviate this issue, we propose to differentiate through the Bellman fixed-point equation to decouple forward and backward passes for Value Iteration Network and its variants, which enables constant backward cost (in planning horizon) and flexible forward budget and helps scale up to large tasks. We study the convergence stability, scalability, and efficiency of the proposed implicit version of VIN and its variants and demonstrate their superiorities on a range of planning tasks: 2D navigation, visual navigation, and 2-DOF manipulation in configuration space and workspace.