Online Residential Demand Response via Contextual Multi-Armed Bandits

Online Residential Demand Response via Contextual Multi-Armed Bandits
复制标题

DOI:
10.1109/lcsys.2020.3003190
复制
发表时间:
2021-04-01
影响因子:
3
通讯作者:
Li, Na
Li, Na
中科院分区:
其他
文献类型:
--
作者:
Chen, Xin;Nie, Yutong;Li, Na

文献摘要

被引文献

相似文献

通过需求响应(DR)计划,住宅负荷具有提高电力系统效率和可靠性的巨大潜力。住宅DR的一个主要挑战是如何学习和处理未知和不确定的客户行为。在这封信中,我们考虑了住宅DR问题,其中负载服务实体(LSE)的目标是选择最优的客户子集来优化某些DR性能,例如最大化具有财务预算的预期负载减少或最小化与目标减少水平的期望平方偏差。为了学习受各种时变环境因素影响的不确定客户行为,我们将住宅DR描述为上下文多臂盗贼(MAB)问题,并提出了一种基于Thompson抽样的在线学习和选择(OLS)算法来解决该问题。该算法考虑了上下文信息,适用于复杂的灾难恢复环境。数值仿真结果表明了该算法的学习有效性。
Residential loads have great potential to enhance the efficiency and reliability of electricity systems via demand response (DR) programs. One major challenge in residential DR is how to learn and handle unknown and uncertain customer behaviors. In this letter, we consider the residential DR problem where the load service entity (LSE) aims to select an optimal subset of customers to optimize some DR performance, such as maximizing the expected load reduction with a financial budget or minimizing the expected squared deviation from a target reduction level. To learn the uncertain customer behaviors influenced by various time-varying environmental factors, we formulate the residential DR as a contextual multi-armed bandit (MAB) problem, and develop an online learning and selection (OLS) algorithm based on Thompson sampling to solve it. This algorithm takes the contextual information into consideration and is applicable to complicated DR settings. Numerical simulations are performed to demonstrate the learning effectiveness of the proposed algorithm.