Perspective view of autonomous control in unknown environment: Dual control for exploitation and exploration vs reinforcement learning

Perspective view of autonomous control in unknown environment: Dual control for exploitation and exploration vs reinforcement learning
复制标题

DOI:
10.1016/j.neucom.2022.04.131
复制
发表时间:
2022-05
期刊:
影响因子:
6
通讯作者:
Wen‐Hua Chen
Wen‐Hua Chen
中科院分区:
计算机科学2区
文献类型:
--
作者:
Wen‐Hua Chen

文献摘要

相似文献

本文综述并讨论了强化学习(RL)和最近发展的开发和探索双重控制(DCEE)之间的关系。有人认为,有两个相关的,但相当独特的方法,即控制和机器学习,在解决棘手的最佳决策/控制问题。在控制方法中,原始问题(无限时域)近似为有限时域问题,并利用计算能力在线求解。在机器学习方法中,最佳解决方案通过迭代来近似,或者当模型不可用时通过试验进行(离线)训练。在处理未知环境时,DCEE作为一种从控制方法发展起来的技术,可以解决与RL类似的问题,同时提供许多优势,最值得注意的是,应对环境/任务中的不确定性,通过平衡开发和探索来提高学习效率,以及建立其稳定性等形式属性的潜力。讨论了DCEE与其他相关方法如双重控制、模型预测控制以及神经科学中的主动推理之间的联系。后者为DCEE提供了强有力的生物学支持。通过使用机器人进行自主源搜索来说明这些方法和讨论。它的结论是,DCEE提供了一个有前途的,补充的方法RL,需要更多的研究来发展它作为一个通用的理论,并充分发挥其潜力。本文揭示的关系为这些相关方法提供了见解,并促进了控制,机器学习和神经科学之间的交叉受精,以在不确定的环境下开发自主控制。
This paper overviews and discusses the relationship between Reinforcement Learning (RL) and the recently developed Dual Control for Exploitation and Exploration (DCEE). It is argued that there are two related but quite distinctive approaches, namely, control and machine learning, in tackling intractability arising in optimal decision making/control problems. In the control approach, the original problems (of an infinite horizon) are approximated by finite horizon problems and solved online by taking advantage of the availability of computing power. In the machine learning approach, the optimal solutions are approximated through iterations, or (offline) training through trials when models are not available. When dealing with unknown environments, DCEE as a technique developed from the control approach could potentially solve similar problems as RL while offering a number of advantages, most notably, coping with uncertainty in environment/tasks, high efficiency in learning through balancing exploitation and exploration, and potential in establishing its formal properties like stability. The links between DCEE and other relevant methods like dual control, Model Predictive Control and particularly Active Inference in neuroscience are discussed. The latter provides a strong biological endorsement for DCEE. The methods and discussions are illustrated by autonomous source search using a robot. It is concluded that DCEE provides a promising, complementary approach to RL, and more research is required to develop it as a generic theory and fully realise its potential. The relationships revealed in this paper provide insights into these relevant methods and facilitate cross fertilisation between control, machine learning and neuroscience for developing autonomous control under uncertain environments.