CntrlDA: A building energy management control system with real-time adjustments. Application to indoor temperature

CntrlDA: A building energy management control system with real-time adjustments. Application to indoor temperature
复制标题

DOI:
10.1016/j.buildenv.2022.108938
复制
发表时间:
2022-03
影响因子:
7.4
通讯作者:
Alex Dmitrewski;Miguel Molina-Solana;Rossella Arcucci
Alex Dmitrewski;Miguel Molina-Solana;Rossella Arcucci
中科院分区:
工程技术1区
文献类型:
--
作者:
Alex Dmitrewski;Miguel Molina-Solana;Rossella Arcucci

文献摘要

相似文献

基于规则的控制 (RBC) 和模型预测控制 (MPC) 传统上用于控制建筑供暖、通风和空调 (HVAC) 系统。然而,当面临在更大层面上有效控制这些系统时,它们存在缺点。强化学习 (RL) 最近成为一种可行的替代方案,与以前的方法相比,显示出有希望的结果,但在未经训练的情况或突然变化时仍然存在一些困难。CntrlDA 是我们的建议,通过将其与数据同化 (DA)(一种数值天气预报中常用的技术)相结合来改进 RL 公式。我们在建筑模拟环境中进行的一系列实验表明,使用 DA 和外部数据训练 RL 控制代理比仅使用模拟数据训练代理具有更好的性能。含有 DA 的 RL 控制剂比不含 DA 的 RL 控制剂维持温度范围的频率高 15.6%。研究还表明,通过在控制过程中包含 DA 阶段,代理可以更好地处理意外事件(这在现实系统中很常见,尤其是在建筑能源控制场景中)。我们表明,与没有 DA 的系统相比,它保持范围的频率提高了 15.4%,并且没有显着增加资源成本。
Rule-Based Control (RBC) and Model Predictive Control (MPC) have been traditionally used to control building heating, ventilation and air conditioning (HVAC) systems. They, however, present shortcomings when faced with efficiently controlling these systems at a larger level. Reinforcement Learning (RL) has recently emerged as a viable alternative, showing promising results compared to previous methods, but still having some difficulties with untrained situations or sudden changes.CntrlDAis our proposal on improving the RL formulation by coupling it with data assimilation (DA), a technique commonly used in numerical weather prediction. Our battery of experiments, in a building simulation environment, shows that training a RL control agent with DA and external data, leads to better performance than training the agent using only the simulation data. The RL control agent with DA maintains the temperature range 15.6% more often than the RL control agent without DA. It is also shown that by including a DA stage in the control process, the agent better deals with unexpected events (which are common in real-life systems and particularly in building energy control scenarios). We show that it maintains the range 15.4% more often than the system without DA with no significant added cost of resources.