Learning and Selecting the Right Customers for Reliability: A Multi-Armed Bandit Approach
Learning and Selecting the Right Customers for Reliability: A Multi-Armed Bandit Approach
复制标题
DOI:
10.1109/cdc.2018.8619481
复制
发表时间:
2018-12
期刊:
影响因子:
--
通讯作者:
Yingying Li;Qinran Hu;Na Li
中科院分区:
文献类型:
--
作者:
Yingying Li;Qinran Hu;Na Li
In this paper, we consider residential demand response (DR) programs where an aggregator calls upon some residential customers to change their demand so that the total load adjustment is as close to a target value as possible. Major challenges lie in the uncertainty and randomness of the customer behaviors in response to DR signals, and the limited knowledge available to the aggregator of the customers. To learn and select the right customers, we formulate the DR problem as a combinatorial multi-armed bandit (CMAB) problem with a reliability goal. We propose a learning algorithm: CUCB-Avg (Combinatorial Upper Confidence Bound-Average), which utilizes both upper confidence bounds and sample averages to balance the tradeoff between exploration (learning) and exploitation (selecting). We prove that CUCB-Avg achieves $O(\log T)$ regret given a time-invariant target. Simulation results demonstrate that our CUCB-Avg performs significantly better than the classic algorithm CUCB (Combinatorial Upper Confidence Bound).