Optimal Control of District Cooling Energy Plant With Reinforcement Learning and MPC

Optimal Control of District Cooling Energy Plant With Reinforcement Learning and MPC
复制标题

利用强化学习和 MPC 的区域供冷能源厂优化控制

DOI:
10.1115/1.4064023
复制
发表时间:
2023
期刊:
ASME Journal of Engineering for Sustainable Buildings and Cities
影响因子:
--
通讯作者:
Barooah, Prabir
Barooah, Prabir
中科院分区:
--
文献类型:
--
作者:
Guo, Zhong;Chaudhari, Aditya;Coffman, Austin R;Barooah, Prabir

文献摘要

相似文献

我们考虑的问题,区域冷却能源工厂(DCEPs)组成的多个冷水机组,冷却塔,和一个热能存储(TES)的最优控制,在随时间变化的电价。由于冷水机组的开/关和DCEP模型的复杂性,模型预测控制(MPC)的直接应用需要求解具有挑战性的混合整数非线性规划(MINLP)。强化学习(RL)是一个有吸引力的替代方案,因为它的实时控制计算要简单得多。但是,由于无数的设计选择和计算密集型训练,设计RL控制器是具有挑战性的。在本文中,我们提出了RL控制器和MPC控制器的DCEP的电力成本最小化,并通过仿真比较。这两个控制器被设计为在目标和信息要求方面具有可比性。RL控制器使用一种新的Q学习算法,该算法是基于最小二乘策略迭代。我们描述了RL控制器的设计选择,包括状态空间和基函数的选择,被发现是有效的。建议MPC控制器不需要一个混合整数求解器的实施,但只有一个非线性规划(NLP)求解器。一个基于规则的基线控制器也提出了援助比较。仿真结果表明,所提出的RL和MPC控制器实现了类似的节省基线控制器,约17%。
We consider the problem of optimal control of district cooling energy plants (DCEPs) consisting of multiple chillers, a cooling tower, and a thermal energy storage (TES), in the presence of time-varying electricity prices. A straightforward application of model predictive control (MPC) requires solving a challenging mixed-integer nonlinear program (MINLP) because of the on/off of chillers and the complexity of the DCEP model. Reinforcement learning (RL) is an attractive alternative since its real-time control computation is much simpler. But designing an RL controller is challenging due to myriad design choices and computationally intensive training. In this paper, we propose an RL controller and an MPC controller for minimizing the electricity cost of a DCEP, and compare them via simulations. The two controllers are designed to be comparable in terms of objective and information requirements. The RL controller uses a novel Q-learning algorithm that is based on least-squares policy iteration. We describe the design choices for the RL controller, including the choice of state space and basis functions, that are found to be effective. The proposed MPC controller does not need a mixed-integer solver for implementation, but only a nonlinear program (NLP) solver. A rule-based baseline controller is also proposed to aid in comparison. Simulation results show that the proposed RL and MPC controllers achieve similar savings over the baseline controller, about 17%.