Accelerate online reinforcement learning for building HVAC control with heterogeneous expert guidances

Accelerate online reinforcement learning for building HVAC control with heterogeneous expert guidances
复制标题

DOI:
10.1145/3563357.3564064
复制
发表时间:
2022-11
期刊:
Proceedings of the 9th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation
影响因子:
--
通讯作者:
Shichao Xu;Yangyang Fu;Yixuan Wang;Zhuoran Yang;Zheng O’Neill;Zhaoran Wang;Qi Zhu
Shichao Xu;Yangyang Fu;Yixuan Wang;Zhuoran Yang;Zheng O’Neill;Zhaoran Wang;Qi Zhu
中科院分区:
其他
文献类型:
--
作者:
Shichao Xu;Yangyang Fu;Yixuan Wang;Zhuoran Yang;Zheng O’Neill;Zhaoran Wang;Qi Zhu

文献摘要

相似文献

建筑供暖、通风和空调(HVAC)系统占美国建筑能耗的近一半以及总能耗的20%。它们的运行对于确保建筑使用者的身心健康也至关重要。与传统的基于模型的HVAC控制方法相比,近期基于无模型深度强化学习(DRL)的方法表现出良好性能,同时不需要开发详细且昂贵的物理模型。然而,这些无模型的DRL方法往往需要很长的训练时间才能达到良好性能,这是它们实际应用的一个主要障碍。在这项工作中,我们提出一种系统方法,通过充分利用来自不同形式的领域专家的知识来加速HVAC控制的在线强化学习。具体而言,算法阶段包括从现有的抽象物理模型以及通过离线强化学习从历史数据中学习专家函数,将专家函数与基于规则的准则相结合,在集成的专家函数指导下进行训练以及从提炼的专家函数进行策略初始化。实验结果表明,与之前基于DRL的方法相比,速度提高了8.8倍。
Building heating, ventilation, and air conditioning (HVAC) systems account for nearly half of building energy consumption and 20% of total energy consumption in the US. Their operation is also crucial for ensuring the physical and mental health of building occupants. Compared with traditional model-based HVAC control methods, the recent model-free deep reinforcement learning (DRL) based methods have shown good performance while do not require the development of detailed and costly physical models. However, these model-free DRL approaches often suffer from long training time to reach a good performance, which is a major obstacle for their practical deployment. In this work, we present a systematic approach to accelerate online reinforcement learning for HVAC control by taking full advantage of the knowledge from domain experts in various forms. Specifically, the algorithm stages include learning expert functions from existing abstract physical models and from historical data via offline reinforcement learning, integrating the expert functions with rule-based guidelines, conducting training guided by the integrated expert function and performing policy initialization from distilled expert function. Experimental results demonstrate up to 8.8X speedup over previous DRL-based methods.