风险敏感的马氏决策过程与强化学习及其应用
批准号:
62073346
项目类别:
面上项目
资助金额:
57.0 万元
负责人:
夏俐
依托单位:
学科分类:
控制理论与技术
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
夏俐
中文摘要
本课题旨在基于马氏决策过程(MDP)和强化学习的理论框架,研究风险敏感的学习优化理论。该类问题的系统累积回报值本质是一个随机变量,传统MDP理论只优化其数学期望,忽视了方差和概率分布信息,本课题主要研究考虑方差约束的多目标MDP优化问题以及考虑概率准则的MDP优化问题,并在此基础上研究数据驱动的强化学习优化算法来有效降低系统收益的风险,完善MDP和强化学习在风险相关性能指标下的理论体系;进一步将上述方法应用于研究金融工程中的资产组合风险管理优化问题和能源互联网中的风储联合出力波动性优化问题,通过方差和概率指标来刻画金融风险和电网安全。本课题的主要学术挑战是处理现有动态决策优化理论在优化风险指标时所面临的本质困难:风险指标的非线性和不可加性使得决策问题不符合标准MDP模型,经典动态规划的一致性选择原则和Bellman最优性方程不再成立。本课题是一个深度交叉的研究前沿,有望取得较好的研究成果。
英文摘要
This proposal aims to study the risk-sensitive stochastic learning and optimization theory, based on the fundamental framework of Markov decision processes (MDPs) and reinforcement learning. The accumulated reward of such stochastic dynamic systems is a random variable. However, many traditional theories only focus on the optimization of its expectation, while ignoring the variance and probabilities. We will first study the variance-related multi-objective MDPs and the probability criterion for MDPs. Based on these risk-sensitive MDP theory, we will further study to develop data-driven reinforcement learning algorithms to efficiently reduce the risk related criterion of stochastic systems, such that a comprehensive framework for risk-sensitive MDP theory and reinforcement learning can be built up. Furthermore, we will apply the risk-sensitive learning and optimization theory to solve the practical problems in financial engineering and energy Internet, including the portfolio management problem for asset returns and risks, and the fluctuation reduction of power output of the renewables and energy storage systems. We use the variance criterion and the probability criterion to quantify the metrics of financial risk and grid safety. The main challenge of the classical dynamic programming theory when applied to risk-related criteria is that the risk-related function is usually nonlinear and non-additive, which induces that the optimization problem does not fit a standard MDP model. The principle of consistent choice of dynamic programming and the Bellman’s optimality equation do not hold any more for such risk-aware MDP problems. Therefore, the research on stochastic optimization problems under risk-related criteria is a young and promising multi-disciplinary area that can potentially produce fruitful research results.
本项目专注于研究风险敏感的随机动态决策优化问题,特别是考虑方差指标和概率指标的MDP(马尔可夫决策过程)优化理论,以及相应的风险敏感强化学习算法,并最终将这些理论应用于能源、金融等工程领域。本项目已达成预期的研究目标,构建了一套风险敏感MDP的优化理论体系。具体成果包括:.1. 提出了最大化均值-方差联合准则的折扣MDP优化方法。.2. 提出了最大化均值-方差联合准则的平均MDP优化方法。.3. 提出了最小化概率准则的折扣MDP优化方法。.4. 提出了最小化概率准则的有限阶段部分可观MDP优化方法。.5. 设计了优化MDP均值-方差及均值-半方差指标的强化学习算法。.这些优化理论和方法已被成功应用于新能源领域的风储联合出力波动性优化问题,以及金融工程中的投资组合管理问题,并取得了显著的应用效果。..总体而言,本项目以方法论研究为核心,同时具有强烈的应用导向。所构建的风险敏感MDP优化理论与算法体系具有广泛的应用前景,可应用于能源、交通、通信、金融等相关领域的一系列涉及风险和安全指标的动态优化决策问题。项目的整体执行情况良好,不仅完成了原计划的研究内容,还进一步深入研究了风险敏感的随机博弈问题,并取得了良好的成果。基于本项目的研究成果,进一步探讨了CVaR(条件风险价值)风险准则的MDP优化方法,并于2023年成功培育获得了新的国家自然科学基金专项项目和面上项目各一项。这些新项目分别重点研究风险相关的新型能源系统协同调度优化和CVaR的动态学习优化理论与算法。
考虑CVaR风险指标的马氏决策过程和强化学习及其应用
-
批准号:72371253
-
项目类别:面上项目
-
资助金额:41万元
-
批准年份:2023
-
负责人:夏俐
-
依托单位:
均值方差动态学习优化理论及其应用
-
批准号:--
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2021
-
负责人:夏俐
-
依托单位:
基于公平指标的排队系统优化理论及应用
-
批准号:61573206
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:夏俐
-
依托单位:
国内基金
海外基金