Progressive research on the exploitation-oriented learning XoL
Progressive research on the exploitation-oriented learning XoL
批准号:
22500143
负责人:
MIYAZAKI Kazuteru
金额:
$2.5万
依托单位国家:
日本
项目类别:
Grant-in-Aid for Scientific Research (C)
财政年份:
2010
资助国家:
日本
项目状态:
已结题
起止时间:
2010 至 2012
中文摘要
本研究完成了一种可处理多种奖励和惩罚的开发导向学习(XoL)方法。此外,XoL方法的奖励和惩罚的设计准则已通过说明性的例子,即课程分类任务,肌腱驱动的机器人的腰部轨迹学习任务,并在多智能体环境中的Keepaway任务提出。指出XoL在应用于实际问题时优于传统的基于动态规划的强化学习。
英文摘要
This research has completed an Exploitation-oriented Learning (XoL) method that can treat multiple rewards and penalties. Furthermore the design guideline of rewards and penalties on the XoL method has been proposed through illustrative examples, namely, a course classification task, a waist-trajectory learning task for a tendon-driven biped robot, and a Keepaway task in a multi-agent environment. It claim that XoL surpass traditional Reinforcement Learning based on Dynamic Programming in application to real-world problem.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Proposal and Evaluation of the Active Course Classification Support System with Exploitation-oriented Learning
探究式学习主动课程分类支持系统的提出与评估
DOI:
10.1007/978-3-642-29946-9_32
发表时间:
2012
期刊:
Lecture Notes in Computer Science
影响因子:
--
作者:
[Seiya Kuroda, Kazuteru Miyazaki and Hiroaki Kobayashi, Kazuteru Miyazaki, Kazuteru Miyazaki and Masaaki Ida]
通讯作者:
Kazuteru Miyazaki and Masaaki Ida
マルチエージェント連続タスクへの改良型罰回避政策形成アルゴリズムの適用とサッカーロボットを用いた実験による評価
改进的惩罚避免策略制定算法在多智能体连续任务中的应用以及通过使用足球机器人的实验进行评估
DOI:
--
发表时间:
2010
期刊:
第53回自動制御連合講演会論文集
影响因子:
--
作者:
[伊藤昌樹, 宮崎和光, 小林博明]
通讯作者:
小林博明
Introduction of Fixed Mode States into Online Reinforcement Learning with Penalty and Reward and Its Application to Waist Trajectory Generation of Biped Robot
将固定模式状态引入带惩罚和奖励的在线强化学习及其在双足机器人腰部轨迹生成中的应用
DOI:
--
发表时间:
2012
期刊:
Journal of Advanced Computational Intelligence and Intelligent Informatics
影响因子:
0.7
作者:
[Seiya Kuroda, Kazuteru Miyazaki and Hiroaki Kobayashi]
通讯作者:
Kazuteru Miyazaki and Hiroaki Kobayashi
複数種類の報酬と罰に対応した経験強化型学習の提案と設計指針に関する研究
研究适应多种奖励和惩罚的体验强化学习的建议和设计指南
DOI:
--
发表时间:
2012
期刊:
平成24年 電気学会 電子・情報・システム部門大会 講演論文集
影响因子:
--
作者:
[Seiya Kuroda, Kazuteru Miyazaki and Hiroaki Kobayashi, Kazuteru Miyazaki, Kazuteru Miyazaki and Masaaki Ida, 宮崎和光]
通讯作者:
宮崎和光
The Penalty Avoiding Rational Policy Making algorithm in Continuous Action Spaces
连续行动空间中避免惩罚的理性决策算法
DOI:
--
发表时间:
2010
期刊:
Proceedings of the 11th International Conference on Intelligent Data Engineering and Automated Learning
影响因子:
--
作者:
[Miyazaki, K.]
通讯作者:
K.
共 21 条
Research on new machine learning method combining Exploitation-oriented Learning and Deep Learning
-
批准号:17K00327
-
项目类别:Grant-in-Aid for Scientific Research (C)
-
资助金额:$2.83万
-
财政年份:2017
-
负责人:MIYAZAKI Kazuteru
-
依托单位:
海外基金