Fast Human-in-the-Loop Control for HVAC Systems via Meta-Learning and Model-Based Offline Reinforcement Learning

Fast Human-in-the-Loop Control for HVAC Systems via Meta-Learning and Model-Based Offline Reinforcement Learning
复制标题

DOI:
10.1109/tsusc.2023.3251302
复制
发表时间:
2023-07
影响因子:
3.9
通讯作者:
Liangliang Chen;Fei Meng;Ying Zhang
Liangliang Chen;Fei Meng;Ying Zhang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Liangliang Chen;Fei Meng;Ying Zhang

文献摘要

相似文献

强化学习(RL)方法可用于开发用于加热、通风和空调(HVAC)系统的控制器,该控制器既节省能量又确保高的居住者热舒适水平。然而,现有的工作通常需要在政策上的数据来训练RL代理,并没有考虑乘员的个性化热偏好,这是在现实世界的场景中的限制。针对个性化空调系统设计了一种基于模型的离线强化学习算法。该算法能够在较少的热反馈下快速适应不同的热偏好,有效地保证了高的个性化热舒适水平。首先,我们使用元监督学习算法来训练乘员的热偏好模型。然后,我们训练一个集成神经网络来预测所考虑的区域的热状态。此外,所获得的集成网络可以指示离线数据集所覆盖的状态和动作空间中的区域。通过元测试更新的个性化热偏好模型,基于模型的RL被用来获得最佳的HVAC控制器。由于所提出的算法只需要离线数据集和一些在线热反馈进行训练,它有助于RL算法更实际地部署到HVAC系统。我们使用ASHRAE数据库II来验证元学习算法对不同居住者热偏好建模的有效性和优势。在EnergyPlus环境下的仿真结果表明,与基于模型的基于策略数据聚合的强化学习算法相比,该算法在保证个性化热偏好的同时,功耗仅增加了1.91%。
Reinforcement learning (RL) methods can be used to develop a controller for the heating, ventilation, and air conditioning (HVAC) systems that both saves energy and ensures high occupants’ thermal comfort levels. However, the existing works typically require on-policy data to train an RL agent, and the occupants’ personalized thermal preferences are not considered, which is limited in the real-world scenarios. This paper designs a high-performance model-based offline RL algorithm for personalized HVAC systems. The proposed algorithm can quickly adapt to different occupants’ thermal preferences with a few thermal feedbacks, guaranteeing the high occupants’ personalized thermal comfort levels efficiently. First, we use a meta-supervised learning algorithm to train an occupant's thermal preference model. Then, we train an ensemble neural network to predict the thermal states of the considered zone. In addition, the obtained ensemble networks can indicate the regions in the state and action spaces covered by the offline dataset. With the personalized thermal preference model updated via meta-testing, model-based RL is used to derive the optimal HVAC controller. Since the proposed algorithm only requires offline datasets and a few online thermal feedbacks for training, it contributes to a more practical deployment of the RL algorithm to HVAC systems. We use the ASHRAE database II to verify the effectiveness and advantage of the meta-learning algorithm for modeling different occupants’ thermal preferences. Numerical simulations on the EnergyPlus environment demonstrate that the proposed algorithm can guarantee personalized thermal preferences with a slight increase of power consumption of 1.91% compared with the model-based RL algorithm with on-policy data aggregation.