Average Reward Reinforcement Learning
Average Reward Reinforcement Learning
批准号:
9520243
负责人:
Prasad Tadepalli
金额:
$22.49万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
1995
资助国家:
美国
项目状态:
已结题
起止时间:
1995-09-15 至 1999-10-31
中文摘要
IRI-9520243 Prasad Tadepalli俄勒冈州州立大学$77,772 - 12个月。 平均奖励强化学习 强化学习是由于奖励和惩罚而提高代理的性能。 到目前为止,大多数对这一现象的研究都假设学习是通过最大化来自环境的贴现总回报来进行的。贴现倾向于牺牲更高的长期回报,以换取短期回报。 另一方面,在真实的情况下,使用单位时间的平均奖励往往能更好地提高绩效。 一个这样的应用是自动引导车辆(AGV-用于柔性制造系统中的材料处理任务的机器人)的调度,其中折扣总奖励的优化将影响机器人服务于靠近的机器而忽略那些远离的机器,导致更远的机器停止,并且可能导致整个制造系统的停止。 该项目的目标是使用AGV调度问题作为研究强化学习方法的一种手段,优化平均奖励以避免此类陷阱,从而更好地了解自动平均奖励强化学习下影响性能的因素。
英文摘要
IRI-9520243 Prasad Tadepalli Oregon State University $77,772 - 12 mos. Average Reward Reinforcement Learning Reinforcement learning is the improvement of performance of an agent due to rewards and punishments. Most studies of the phenomenon so far have assumed that the learning proceeds through maximizing the discounted total reward from the environment. Discounting tends to favor the sacrifice of higher long-term rewards in favor of short-term rewards. On the other hand, performance is often better enhanced in real situations by the use of average reward per unit time. One such application is the scheduling of automatic guided vehicles (AGVs - robots used for material handling tasks in flexible manufacturing systems), where optimization of discounted total reward will influence the robots to serve machines that are close and neglect those far away, leading to stoppage of the more distant machines, and perhaps to the halting of the entire manufacturing system. The objective of this project is to use the AGV scheduling problem as a means of studying reinforcement learning methods that optimize average reward to avoid such pitfalls, leading to a better understanding of the factors that influence performance under automated average reward reinforcement learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Integrating Learning and Search for Structured Prediction
-
批准号:1219258
-
项目类别:Standard Grant
-
资助金额:$45.0万
-
财政年份:2012
-
负责人:Prasad Tadepalli
-
依托单位:
RI: Medium: Collaborative Research: Optimizing Policies for Service Organizations in Complex Structured Domains
-
批准号:0964705
-
项目类别:Continuing Grant
-
资助金额:$58.17万
-
财政年份:2010
-
负责人:Prasad Tadepalli
-
依托单位:
Relational Reinforcement Learning
-
批准号:0329278
-
项目类别:Continuing Grant
-
资助金额:$41.16万
-
财政年份:2003
-
负责人:Prasad Tadepalli
-
依托单位:
Average Reward Reinforcement Learning: Scaling up
-
批准号:0098050
-
项目类别:Continuing Grant
-
资助金额:$36.87万
-
财政年份:2001
-
负责人:Prasad Tadepalli
-
依托单位:
Tradeoffs in Learning and Planning
-
批准号:9111231
-
项目类别:Standard Grant
-
资助金额:$7.0万
-
财政年份:1991
-
负责人:Prasad Tadepalli
-
依托单位:
海外基金