CAREER: Soft-robust Methods for Offline Reinforcement Learning
CAREER: Soft-robust Methods for Offline Reinforcement Learning
批准号:
2144601
负责人:
Marek Petrik
金额:
$57.59万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-01 至 2027-08-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Improvements in sensors, data collection, and computational power have driven the desire to harness data to improve decision-making in various domains, such as precision agriculture or medicine. Making effective data-driven decisions is a challenging problem studied by the reinforcement learning community. Despite their success in simulated domains, like board games, existing reinforcement learning algorithms are still too unreliable for widespread deployment. This project develops new reinforcement learning algorithms that achieve reliability by carefully balancing the expected quality of recommended decisions with their risk of failure. The risk of failure is measured using techniques from stochastic finance, which enable flexible and efficient algorithms. These new, reliable algorithms will help bring data-driven decision-making to new domains, including agriculture, ecological management, and medicine. The research is integrated with educational activities to provide graduate and undergraduate students with training opportunities and new study materials, including a textbook. This research develops and analyzes algorithms for reinforcement learning (RL) problems with 1) limited or expensive data and 2) a high cost of failure. Soft-robust objectives build on convex risk measures to balance robustness and average quality of data-driven decision-making problems. While soft-robust objectives are well understood in single-stage optimization problems, many fundamental and practical questions remain open in multi-stage problems, like RL. This project will answer these questions and develop reliable, tractable, and scalable algorithms by tackling three main objectives. First, the project will establish soft-robust RL formulations' statistical and computational properties. Second, the project will generate a new class of tabular soft-robust RL algorithms built on new insights into the relationship between soft-robustness and robust Markov decision processes. Third, the project will scale the tabular algorithms to value function approximation and gradient style methods. The project will address both batch RL, with known rewards and estimated transition probabilities, and inverse RL, with estimated rewards and known transition probabilities.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
Solving multi-model MDPs by coordinate ascent and dynamic programming
通过坐标上升和动态规划求解多模型 MDP
DOI:
--
发表时间:
2023
期刊:
Conference on Uncertainty in Artificial Intelligence
影响因子:
--
作者:
[Xihong Su, Marek Petrik]
通讯作者:
Xihong Su, Marek Petrik
DOI:
--
发表时间:
2023
期刊:
影响因子:
--
作者:
[J. Hau;Marek Petrik;M. Ghavamzadeh]
通讯作者:
J. Hau;Marek Petrik;M. Ghavamzadeh
DOI:
--
发表时间:
2022-12
期刊:
影响因子:
--
作者:
[Qiuhao Wang;C. Ho;Marek Petrik]
通讯作者:
Qiuhao Wang;C. Ho;Marek Petrik
RI: SMALL: Robust Reinforcement Learning Using Bayesian Models
-
批准号:1815275
-
项目类别:Standard Grant
-
资助金额:$43.78万
-
财政年份:2018
-
负责人:Marek Petrik
-
依托单位:
III: Small: Robust Reinforcement Learning for Invasive Species Management
-
批准号:1717368
-
项目类别:Standard Grant
-
资助金额:$49.73万
-
财政年份:2017
-
负责人:Marek Petrik
-
依托单位:
海外基金