Achieving robustness by optimizing failure behavior

Achieving robustness by optimizing failure behavior
复制标题

DOI:
10.1109/icra.2017.7989681
复制
发表时间:
2017-05
期刊:
2017 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Manuel Baum;O. Brock
Manuel Baum;O. Brock
中科院分区:
其他
文献类型:
--
作者:
Manuel Baum;O. Brock

文献摘要

相似文献

学习操作技能的最突出标准是任务成功的最优化,以预期回报或成功概率为模型。如果我们只想优化单个控制器,这是合理的。但是,如果学习的操作原语被用作较大系统中的模块,那么它们生成的传感器跟踪有助于识别动作结果也是很重要的。仅为原语的预期成功进行优化并不能保证这一点。我们演示了一个简单的例子,用于优化面向可观察性的操作,并结合优化以实现预期的成功。我们的实验是一个带有软机械手的操作任务,其中一个动作基元被学习,以便它生成的传感器轨迹有助于分类器区分任务成功和任务失败。实验结果表明,在原有操作基元的基础上增加辅助力确实可以促进操作任务的结果识别。
The most prominent criterion for learning of manipulation skills is the optimization of task success, modeled as expected reward or probability of success. This is sensible if we only want to optimize a single controller. But if learned manipulation primitives are used as modules in a larger system, then it is also important that their generated sensor traces facilitate recognition of action-outcomes. Optimization solely for expected success of a primitive does not guarantee this. We demonstrate a simple example for optimization of actions towards observability, combined with optimization for expected success. Our experiment is a manipulation task with a soft manipulator, where an action primitive is learned such that its generated sensor trace helps a classifier to distinguish task success and task failure. The experimental results indicate that adding auxiliary forces to the original manipulation primitive can indeed facilitate outcome recognition for manipulation tasks.