Long-term Planning by Short-term Prediction

Long-term Planning by Short-term Prediction
复制标题

DOI:
--
复制
发表时间:
2016-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Shai Shalev-Shwartz;Nir Ben-Zrihem;Aviad Cohen;A. Shashua
Shai Shalev-Shwartz;Nir Ben-Zrihem;Aviad Cohen;A. Shashua
中科院分区:
其他
文献类型:
--
作者:
Shai Shalev-Shwartz;Nir Ben-Zrihem;Aviad Cohen;A. Shashua

文献摘要

被引文献

相似文献

我们考虑了在自动驾驶应用中经常出现的规划问题,其中智能体应该决定立即行动以优化长期目标。例如,当一辆汽车试图在环形交叉路口合并时,它应该立即决定加速/制动命令,而该命令的长期影响是合并的成功/失败。这些问题的特点是连续的状态和动作空间,以及与多个代理的交互,这些代理的行为可能是敌对的。我们认为,由于自然状态空间表示的非马尔可夫性,以及连续的状态和动作空间,双重版本的MDP框架(依赖于值函数和$Q$函数)在自动驾驶应用中存在问题。我们建议通过将问题分解为两个阶段来解决规划任务:首先,我们应用监督学习来基于现在预测不久的将来。我们要求预测器相对于现在的表示是可微的。其次,我们使用递归神经网络对智能体的完整轨迹进行建模,其中无法解释的因素被建模为(加性)输入节点。这使我们能够使用监督学习技术和循环神经网络的直接优化来解决长期规划问题。我们的方法使我们能够通过将敌对因素纳入环境中来学习稳健的政策。
We consider planning problems, that often arise in autonomous driving applications, in which an agent should decide on immediate actions so as to optimize a long term objective. For example, when a car tries to merge in a roundabout it should decide on an immediate acceleration/braking command, while the long term effect of the command is the success/failure of the merge. Such problems are characterized by continuous state and action spaces, and by interaction with multiple agents, whose behavior can be adversarial. We argue that dual versions of the MDP framework (that depend on the value function and the $Q$ function) are problematic for autonomous driving applications due to the non Markovian of the natural state space representation, and due to the continuous state and action spaces. We propose to tackle the planning task by decomposing the problem into two phases: First, we apply supervised learning for predicting the near future based on the present. We require that the predictor will be differentiable with respect to the representation of the present. Second, we model a full trajectory of the agent using a recurrent neural network, where unexplained factors are modeled as (additive) input nodes. This allows us to solve the long-term planning problem using supervised learning techniques and direct optimization over the recurrent neural network. Our approach enables us to learn robust policies by incorporating adversarial elements to the environment.