Learning and Selecting the Right Customers for Reliability: A Multi-Armed Bandit Approach

Learning and Selecting the Right Customers for Reliability: A Multi-Armed Bandit Approach
复制标题

DOI:
10.1109/cdc.2018.8619481
复制
发表时间:
2018-12
期刊:
2018 IEEE Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Yingying Li;Qinran Hu;Na Li
Yingying Li;Qinran Hu;Na Li
中科院分区:
其他
文献类型:
--
作者:
Yingying Li;Qinran Hu;Na Li

文献摘要

被引文献

相似文献

在本文中,我们考虑住宅需求响应(DR)规划,其中聚合器呼吁一些住宅客户改变他们的需求,以使总负荷调整尽可能接近目标值。主要的挑战在于客户响应DR信号的行为的不确定性和随机性,以及客户聚集者可用的有限的知识。为了学习和选择合适的客户,我们将DR问题描述为一个以可靠性为目标的组合多臂强盗(CMAB)问题。我们提出了一种学习算法:CUCB-Avg(组合上置信限-平均值),它同时利用上置信限和样本平均来平衡探索(学习)和开发(选择)之间的权衡。我们证明了在给定时不变目标的情况下,CUCB-Avg算法达到了$O(\logT)$RELERY。仿真结果表明,我们的CUCB-Avg算法的性能明显好于经典的组合置信限算法CUCB。
In this paper, we consider residential demand response (DR) programs where an aggregator calls upon some residential customers to change their demand so that the total load adjustment is as close to a target value as possible. Major challenges lie in the uncertainty and randomness of the customer behaviors in response to DR signals, and the limited knowledge available to the aggregator of the customers. To learn and select the right customers, we formulate the DR problem as a combinatorial multi-armed bandit (CMAB) problem with a reliability goal. We propose a learning algorithm: CUCB-Avg (Combinatorial Upper Confidence Bound-Average), which utilizes both upper confidence bounds and sample averages to balance the tradeoff between exploration (learning) and exploitation (selecting). We prove that CUCB-Avg achieves $O(\log T)$ regret given a time-invariant target. Simulation results demonstrate that our CUCB-Avg performs significantly better than the classic algorithm CUCB (Combinatorial Upper Confidence Bound).