Reinforcement Learning for Optimal Control of a District Cooling Energy Plant

Reinforcement Learning for Optimal Control of a District Cooling Energy Plant
复制标题

DOI:
10.23919/acc53348.2022.9867239
复制
发表时间:
2022-03
期刊:
2022 American Control Conference (ACC)
影响因子:
--
通讯作者:
Zhong Guo;Austin R. Coffman;P. Barooah
Zhong Guo;Austin R. Coffman;P. Barooah
中科院分区:
其他
文献类型:
--
作者:
Zhong Guo;Austin R. Coffman;P. Barooah

文献摘要

相似文献

区域冷却能源工厂(DCEP)由冷水机组,冷却塔和热能储存(TES)系统组成,消耗大量的电力。利用时变电价优化TES和冷水机组的调度是一个具有挑战性的最优控制问题。经典的方法,模型预测控制(MPC),需要解决一个高维的混合整数非线性规划(MINLP),因为开/关驱动的冷水机组和充电/放电的TES,这是计算上的挑战。RL是一个有吸引力的替代方案:真实的时间控制计算是一个低维的优化问题,可以很容易地解决。然而,强化学习控制器的性能取决于许多设计选择。在本文中,我们提出了一个Q-学习的强化学习(RL)控制器的这个问题。数值模拟结果表明,建议的RL控制器是能够减少能源成本超过基于规则的基线控制器约8%,与文献中报道的节省与MPC类似的DCEP。我们描述了RL控制器的设计选择,包括基函数,奖励函数成形和学习算法参数。与现有的DCEP强化学习的工作相比,所提出的控制器是针对连续的状态和动作空间设计的。
District cooling energy plants (DCEPs) consisting of chillers, cooling towers, and thermal energy storage (TES) systems consume a considerable amount of electricity. Optimizing the scheduling of the TES and chillers to take advantage of time-varying electricity price is a challenging optimal control problem. The classical method, model predictive control (MPC), requires solving a high dimensional mixed-integer nonlinear program (MINLP) because of the on/off actuation of the chillers and charge/discharge of TES, which are computationally challenging. RL is an attractive alternative: the real time control computation is a low-dimensional optimization problem that can be easily solved. However, the performance of an RL controller depends on many design choices.In this paper, we propose a Q-learning based reinforcement learning (RL) controller for this problem. Numerical simulation results show that the proposed RL controller is able to reduce energy cost over a rule-based baseline controller by approximately 8%, comparable to savings reported in the literature with MPC for similar DCEPs. We describe the design choices in the RL controller, including basis functions, reward function shaping, and learning algorithm parameters. Compared to existing work on RL for DCEPs, the proposed controller is designed for continuous state and actions spaces.