RI: Small: Non-parametric Approximate Dynamic Programming for Continuous Domains
RI: Small: Non-parametric Approximate Dynamic Programming for Continuous Domains
批准号:
1218931
负责人:
Ronald Parr
金额:
$45.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-08-01 至 2018-07-31
中文摘要
这个项目涉及一种被称为强化学习的机器学习技术,它与心理学中使用的强化学习概念有关,但又有所不同。 共同点是,这两种观点都研究由经验引起的行为变化。 在机器学习的情况下,行为通常是动态环境中的决策,例如控制机器人,工厂,仓库的库存水平甚至药物剂量水平。 目前这一领域的理论发展保证了强化学习算法可以做出最佳决策,但只能在限制性假设下做出,而这些假设在实践中很难保证。 将强化学习应用于重大实际问题的努力已经取得了一些成功,但这些努力往往放弃理论保证,并依赖于专家进行繁琐的参数调整(人工试错)来实现成功。这项研究旨在减少使强化学习成功所需的人工试错量,从而使其成为更广泛人群更容易使用的工具。 具体来说,它将专注于由连续变量描述的域的算法,寻求为这些域提供更强的理论保证,以及平衡尝试新事物的预期好处与坚持已知问题的好处(探索与利用)的方法。 在这个领域取得成功的一个实际好处是改进技术,使人们更容易部署学习和提高性能的算法,在各种实际任务中,如上面提到的:机器人或工厂控制,库存管理,或药物交付。这个项目计划使用模型直升机作为挑战域,但它本身不是关于直升机控制。 相反,它寻求开发可以应用于许多问题的通用技术,包括直升机,并将使用模型直升机作为一种廉价而有趣的方式来激励学生。 该项目旨在开发一个模型直升机模拟器(以降低在实际直升机上尝试一切的成本和风险),并计划将该模拟器提供给研究界,提供一个有趣且具有挑战性的基准问题。
英文摘要
This project concerns a machine learning technique known as reinforcement learning, which is related to, but distinct from, the notion of reinforcement learning used in psychology. The common element is that both views study changes in behavior that result from experience. In the machine learning case, the behaviors are often decision making in dynamic environments, such as controlling a robot, a factory, inventory levels for a warehouse or even drug dosage levels. Current theoretical development in this area guarantees that optimal decisions can be made by reinforcement learning algorithms, but only under restrictive assumptions that are difficult to ensure in practice. Efforts to apply reinforcement learning to significant practical problems have enjoyed some success, but such efforts often forgo theoretical guarantees and rely upon tedious parameter adjustments by experts (human trial and error) to achieve success.This research seeks to reduce the amount of human trial and error needed to make reinforcement learning successful, thereby making it a more accessible tool to a wider range of people. Specifically, it will focus on algorithms for domains described by continuous variables, seeking to provide stronger theoretical guarantees for such domains as well as an approach that balances the anticipated benefit of trying new things with the benefit of sticking to what is already known about a problem (exploration vs. exploitation). A practical benefit of success in this area would be improved techniques that make it easier for people to deploy algorithms that learn and improve performance in a variety of practical tasks like those mentioned above: robot or factory control, inventory management, or drug delivery.This project plans to use a model helicopter as a challenge domain, but it is not about helicopter control per se. Rather, it seeks to develop general techniques that can apply to many problems, including helicopters, and will use model helicopters as an inexpensive and fun way to motivate students. The project aims to develop a model helicopter simulator (to reduce the cost and risk of trying everything on an actual helicopter) and plans to make this simulator available to the research community, providing a fun and challenging benchmark problem.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Feature Encoding for Reinforcement Learning
-
批准号:1815300
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2018
-
负责人:Ronald Parr
-
依托单位:
EAGER: Collaborative Research: An Unified Learnable Roadmap for Sequential Decision Making in Relational Domains
-
批准号:1836575
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2018
-
负责人:Ronald Parr
-
依托单位:
EAGER: IIS: RI: Learning in Continuous and High Dimensional Action Spaces
-
批准号:1147641
-
项目类别:Standard Grant
-
资助金额:$15.0万
-
财政年份:2011
-
负责人:Ronald Parr
-
依托单位:
Collaborative: RI: Feature Discovery and Benchmarks for Exportable Reinforcement Learning
-
批准号:0713435
-
项目类别:Standard Grant
-
资助金额:$22.5万
-
财政年份:2007
-
负责人:Ronald Parr
-
依托单位:
CAREER: Observing to Plan - Planning to Observe
-
批准号:0546709
-
项目类别:Continuing Grant
-
资助金额:$44.0万
-
财政年份:2006
-
负责人:Ronald Parr
-
依托单位:
Prediction and Planning: Bridging the Gap
-
批准号:0209088
-
项目类别:Standard Grant
-
资助金额:$29.17万
-
财政年份:2002
-
负责人:Ronald Parr
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: