Towards Optimal Pricing of Demand Response - A Nonparametric Constrained Policy Optimization Approach

Towards Optimal Pricing of Demand Response - A Nonparametric Constrained Policy Optimization Approach
复制标题

DOI:
10.1109/pesgm52003.2023.10252418
复制
发表时间:
2023-06
期刊:
2023 IEEE Power & Energy Society General Meeting (PESGM)
影响因子:
--
通讯作者:
Jun Song;Chaoyue Zhao
Jun Song;Chaoyue Zhao
中科院分区:
其他
文献类型:
--
作者:
Jun Song;Chaoyue Zhao

文献摘要

相似文献

需求响应(DR)已被证明是减少高峰负荷和减轻电力市场供需双方不确定性的有效方法。减灾研究的一个关键问题是如何适当调整电价,将电力负荷从高峰转移到低峰。近年来,强化学习(RL)已被用来解决基于价格的灾难恢复问题,因为它是一种无模型技术,不需要为最终用户识别模型。然而,大多数强化学习方法无法保证学习定价策略的稳定性和最优性,这在安全关键型电力系统中是不可取的,并可能导致高昂的客户账单。在本文中,我们提出了一种创新的非参数约束策略优化方法,通过消除大多数 RL 文献采用的策略表示的限制性假设:策略必须参数化或落入某个分布类别,在提高最优性的同时确保策略更新的稳定性。我们推导出每次迭代的最优策略更新的封闭式表达式,并开发一种有效的策略行为者批评算法来解决所提出的约束策略优化问题。两个 DR 案例的实验表明,与最先进的 RL 算法相比,我们提出的非参数约束策略优化方法具有优越的性能。
Demand response (DR) has been demonstrated to be an effective method for reducing peak load and mitigating uncertainties on both the supply and demand sides of the electricity market. One critical question for DR research is how to appropriately adjust electricity prices in order to shift electrical load from peak to off-peak hours. In recent years, reinforcement learning (RL) has been used to address the price-based DR problem because it is a model-free technique that does not necessitate the identification of models for end-use customers. However, the majority of RL methods cannot guarantee the stability and optimality of the learned pricing policy, which is undesirable in safety-critical power systems and may result in high customer bills. In this paper, we propose an innovative nonparametric constrained policy optimization approach that improves optimality while ensuring stability of the policy update, by removing the restrictive assumption on policy representation that the majority of the RL literature adopts: the policy must be parameterized or fall into a certain distribution class. We derive a closed-form expression of optimal policy update for each iteration and develop an efficient on-policy actor-critic algorithm to address the proposed constrained policy optimization problem. The experiments on two DR cases show the superior performance of our proposed nonparametric constrained policy optimization method compared with state-of-the-art RL algorithms.