Residential Demand Response of Thermostatically Controlled Loads Using Batch Reinforcement Learning

Residential Demand Response of Thermostatically Controlled Loads Using Batch Reinforcement Learning
复制标题

DOI:
10.1109/tsg.2016.2517211
复制
发表时间:
2017-09
影响因子:
9.6
通讯作者:
F. Ruelens;B. Claessens;Stijn Vandael;B. Schutter;Robert Babuška;R. Belmans
F. Ruelens;B. Claessens;Stijn Vandael;B. Schutter;Robert Babuška;R. Belmans
中科院分区:
工程技术1区
文献类型:
--
作者:
F. Ruelens;B. Claessens;Stijn Vandael;B. Schutter;Robert Babuška;R. Belmans

文献摘要

被引文献

相似文献

在批量强化学习(RL)的最新进展的驱动下,本文有助于批量RL的需求响应的应用。与传统的基于模型的方法相比,批量RL技术不需要系统识别步骤,使其更适合大规模实现。本文扩展了拟合Q迭代,一个标准的批量RL技术,当预测的外生数据提供的情况下。一般来说,批量RL技术不依赖于关于系统动力学或解决方案的专家知识。然而,如果提供了一些专家知识,则可以通过使用所提出的政策调整方法来并入。最后,我们解决了找到参与日前市场所需的开环时间表的挑战。我们提出了一个无模型的蒙特卡罗方法,使用一个度量的基础上的状态-动作值函数或Q-函数,我们说明了这种方法,找到一个热泵恒温器的前一天的时间表。我们的实验表明,批量强化学习技术为基于模型的控制器提供了一种有价值的替代方案,并且它们可用于构建闭环和开环策略。
Driven by recent advances in batch Reinforcement Learning (RL), this paper contributes to the application of batch RL to demand response. In contrast to conventional model-based approaches, batch RL techniques do not require a system identification step, making them more suitable for a large-scale implementation. This paper extends fitted Q-iteration, a standard batch RL technique, to the situation when a forecast of the exogenous data is provided. In general, batch RL techniques do not rely on expert knowledge about the system dynamics or the solution. However, if some expert knowledge is provided, it can be incorporated by using the proposed policy adjustment method. Finally, we tackle the challenge of finding an open-loop schedule required to participate in the day-ahead market. We propose a model-free Monte Carlo method that uses a metric based on the state-action value function or Q-function and we illustrate this method by finding the day-ahead schedule of a heat-pump thermostat. Our experiments show that batch RL techniques provide a valuable alternative to model-based controllers and that they can be used to construct both closed-loop and open-loop policies.