Reinforcement learning of occupant behavior model for cross-building transfer learning to various HVAC control systems

Reinforcement learning of occupant behavior model for cross-building transfer learning to various HVAC control systems
复制标题

DOI:
10.1016/j.enbuild.2021.110860
复制
发表时间:
2021-03-10
影响因子:
6.7
通讯作者:
Chen, Qingyan
Chen, Qingyan
中科院分区:
工程技术2区
文献类型:
--
作者:
Deng, Zhipeng;Chen, Qingyan

文献摘要

被引文献

相似文献

居住者行为在建筑物性能评价中起着重要的作用。然而,许多环境因素,如占用,机械系统和室内设计,对居住者的行为有显着的影响。以往的研究大多建立了数据驱动的行为模型,其可扩展性和泛化能力有限。我们的研究建立了一个基于策略的强化学习(RL)模型,用于调节恒温器和服装水平的行为。乘客行为被建模为马尔可夫决策过程(MDP)。MDP中的动作和状态空间包含乘员行为和各种碰撞参数。居住者行为的目标是一个更舒适的环境,我们模拟了调整行动的奖励,作为行动前后的热感觉投票(TSV)的绝对差异。我们使用Q学习在MATLAB中训练RL模型,并使用收集的数据验证模型。经过训练,该模型预测了将恒温器设定点调整为R-2从0.75到0.8的行为,并且在办公楼中的平均绝对误差(MAE)小于1.1摄氏度(2华氏度)。本研究也将RL模型的行为知识转移到其他具有不同HVAC控制系统的办公建筑中。迁移学习模型预测乘员行为的R-2从0.73到0.8,并且大多数时间MAE小于1.1摄氏度(2华氏度)。从办公楼到住宅楼,迁移学习模型的R-2也超过了0.6。因此,结合迁移学习的强化学习模型能够准确地预测建筑物使用者的行为,具有良好的可扩展性,并且不需要数据收集。(C)2021爱思唯尔有限公司版权所有。
Occupant behavior plays an important role in the evaluation of building performance. However, many contextual factors, such as occupancy, mechanical system and interior design, have a significant impact on occupant behavior. Most previous studies have built data-driven behavior models, which have limited scalability and generalization capability. Our investigation built a policy-based reinforcement learning (RL) model for the behavior of adjusting the thermostat and clothing level. Occupant behavior was modelled as a Markov decision process (MDP). The action and state space in the MDP contained occupant behavior and various impact parameters. The goal of the occupant behavior was a more comfortable environment, and we modelled the reward for the adjustment action as the absolute difference in the thermal sensation vote (TSV) before and after the action. We used Q-learning to train the RL model in MATLAB and validated the model with collected data. After training, the model predicted the behavior of adjusting the thermostat set point with R-2 from 0.75 to 0.8, and the mean absolute error (MAE) was less than 1.1 degrees C (2 degrees F) in an office building. This study also transferred the behavior knowledge of the RL model to other office buildings with different HVAC control systems. The transfer learning model predicted the occupant behavior with R-2 from 0.73 to 0.8, and the MAE was less than 1.1 degrees C (2 degrees F) most of the time. Going from office buildings to residential buildings, the transfer learning model also had an R-2 over 0.6. Therefore, the RL model combined with transfer learning was able to predict the building occupant behavior accurately with good scalability, and without the need for data collection. (C) 2021 Elsevier B.V. All rights reserved.