课题基金 / 基金详情

application of scalable safe reinforcement learning to high-risk robotics

application of scalable safe reinforcement learning to high-risk robotics
可扩展安全强化学习在高风险机器人技术中的应用
批准号:
21J15633
负责人:
Zhu Lingwei
金额:
$0.96万
依托单位国家:
日本
项目类别:
Grant-in-Aid for JSPS Fellows
财政年份:
2021
资助国家:
日本
项目状态:
已结题
起止时间:
2021-04-28 至 2023-03-31

项目摘要

项目成果

相关文献

中文摘要
翻译
总之,研究进展顺利。强化学习高风险管控成果按计划迈出重要步伐。正如研究计划中所述,第一年侧重于解决理论问题。研究结果解决了如何利用熵进行更鲁棒的强化学习框架和后续风险敏感控制的根本问题。这些工作试图从几个不同的角度来解决这个问题,例如直接增加鲁棒性;确保学习的改进;以及利用更稳定的Tsallis熵。因此,以下论文已经发表或向顶级国际会议提交了5篇论文:[1]谨慎的演员评论家,2021年亚洲机器学习会议;[2] Geometric Value Iteration - Dynamic Error Aware KL Regularization for Reinforcement Learning,2021年亚洲机器学习会议; [3] q-Munchausen Reinforcement Learning,2022年人工智能中的不确定性(正在审查中); [4]通过优势学习在最大Tsallis熵框架中执行KL正则化,人工智能中的不确定性2022(审查中); [5]下限最大化单调政策改进,人工智能的不确定性2022(审查中)
英文摘要
As a summary the progress of research has been going well. Important steps towards the achievements of reinforcement learning for high-risk control have been made as planned. As stated in the research plan, the first year focuses on solving the theoretical problems. The results solved the fundamental problem of how to make use of entropy for more robust reinforcement learning framework and subsequent risk-sensitive control. The works attempted to tackle the problem from several different perspectives such as increasing the robustness directly; ensuring learning improvement; and making use of more stable Tsallis entropy.As a result, the following papers have been published/submitted 5 papers to top international conferences: [1] Cautious Actor Critic, Asian Conference on Machine Learning 2021; [2] Geometric Value Iteration - Dynamic Error Aware KL Regularization for Reinforcement Learning, Asian Conference on Machine Learning 2021; [3] q-Munchausen Reinforcement Learning, Uncertainty in Artificial Intelligence 2022 (under review); [4] Enforcing KL Regularization in Maximum Tsallis Entropy Framework via Advantage Learning, Uncertainty in Artificial Intelligence 2022 (under review); [5] Lower Bound Maximizing Monotonic Policy Improvement, Uncertainty in Artificial Intelligence 2022 (under review)
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2021-07
期刊: ArXiv
影响因子: --
作者: [Lingwei Zhu;Toshinori Kitamura;Takamitsu Matsubara]
通讯作者: Lingwei Zhu;Toshinori Kitamura;Takamitsu Matsubara