Risk-Sensitive Learning and Pricing for Demand Response

Risk-Sensitive Learning and Pricing for Demand Response
复制标题

DOI:
10.1109/tsg.2017.2700458
复制
发表时间:
2016-11
影响因子:
9.6
通讯作者:
Kia Khezeli;E. Bitar
Kia Khezeli;E. Bitar
中科院分区:
工程技术1区
文献类型:
--
作者:
Kia Khezeli;E. Bitar

文献摘要

被引文献

相似文献

我们考虑这样一种情况:电力公司试图通过向一组固定客户提供相对于其预定基准的消耗量减少的统一价格来减少其峰值电力需求。基本需求曲线描述了响应于所提供价格的消费总量减少,假设是仿射的并且受到不可观察的随机冲击。假设需求曲线的参数和随机冲击的分布最初对于公用事业公司来说都是未知的,我们研究公用事业公司可以在多大程度上动态调整其报价,以在有限的 ${T}$ 天数内最大化其累积风险敏感收益。为了有效地做到这一点,公用事业公司必须设计其定价政策,以平衡学习未知需求模型(探索)的需要与随着时间的推移最大化其回报(利用)之间的权衡。在本文中,我们提出了这样的定价策略,相对于了解底层需求模型的预言机定价策略,该定价策略在 ${T}$ 天内的预期收益损失最多为 ${O(\sqrt {T}\log (T))}$ 。此外,所提出的定价政策被证明会产生一系列在均方意义上收敛于预言机最优价格的价格。
We consider the setting in which an electric power utility seeks to curtail its peak electricity demand by offering a fixed group of customers a uniform price for reductions in consumption relative to their predetermined baselines. The underlying demand curve, which describes the aggregate reduction in consumption in response to the offered price, is assumed to be affine and subject to unobservable random shocks. Assuming that both the parameters of the demand curve and the distribution of the random shocks are initially unknown to the utility, we investigate the extent to which the utility might dynamically adjust its offered prices to maximize its cumulative risk-sensitive payoff over a finite number of ${T}$ days. In order to do so effectively, the utility must design its pricing policy to balance the tradeoff between the need to learn the unknown demand model (exploration) and maximize its payoff (exploitation) over time. In this paper, we propose such a pricing policy, which is shown to exhibit an expected payoff loss over ${T}$ days that is at most ${O(\sqrt {T}\log (T))}$ , relative to an oracle pricing policy that knows the underlying demand model. Moreover, the proposed pricing policy is shown to yield a sequence of prices that converge to the oracle optimal prices in the mean square sense.