Performance, robustness, and portability of imitation-assisted reinforcement learning policies for shading and natural ventilation control

Performance, robustness, and portability of imitation-assisted reinforcement learning policies for shading and natural ventilation control
复制标题

DOI:
10.1016/j.apenergy.2023.121364
复制
发表时间:
2023-10
期刊:
影响因子:
11.2
通讯作者:
Bumsoo Park;A. Rempel;Sandipan Mishra
Bumsoo Park;A. Rempel;Sandipan Mishra
中科院分区:
工程技术1区
文献类型:
--
作者:
Bumsoo Park;A. Rempel;Sandipan Mishra

文献摘要

相似文献

摘要空间加热和冷却约占所有建筑相关能源消耗的一半,每年排放3 Gt的CO2,或近10%的全球总量。可操作的遮阳、自然通风和太阳能加热是减少这些排放的有希望的策略,利用最小的机械能来调节凉爽的夜间空气、寒冷的夜空和太阳辐射。然而,这些战略没有得到充分利用,因为它们的绩效取决于其可操作要素之间的严格协调。此外,这种系统的个性,以及缺乏适合控制设计的基于物理的模型,阻碍了广泛适用的控制策略的发展。为了解决这个问题,我们在这里开发了一种新的数据驱动策略,用于使用基于策略的强化学习(RL)设计住宅建筑的遮阳和自然通风控制。为了限制不良行为并减少训练时间,我们首先使用模仿学习来初始化具有专家知识的RL训练,产生了一个初始策略,在24个气候多样的城市中将模拟的晚春空间调节负荷减少了≥ 40%。然后,在代表地中海、半干旱、湿润亚热带和大陆性气候的四个城市,用RL对这一政策进行了培训。当部署在不熟悉但相关气候的城市时,这些新政策在潮湿的亚热带地区将空间空调负荷降低了≥ 50%,在其他三种气候中降低了≥ 90%,显示出卓越的便携性。此外,它们的性能对住宅取向、玻璃、内部热增益和空气泄漏的变化出乎意料地稳健。这些结果表明,模仿辅助RL在开发动态被动加热和冷却控制的高性能政策方面具有非凡的潜力,这些政策在不熟悉的情况下仍然有效,消除了无碳建筑运营中被动系统进步的重大障碍。
Abstract Space heating and cooling account for approximately half of all building-related energy consumption, emitting 3 Gt of CO 2 annually, or nearly 10% of the global total. Operable shading, natural ventilation, and solar heating are promising strategies for reducing these emissions, leveraging minimal mechanical energy to condition space with cool night air, cold night skies, and solar radiation. However, these strategies are under-utilized because their performance depends on rigorous coordination among their operable elements. Additionally, the individuality of such systems, and the lack of physics-based models suitable for control design, have thwarted the development of widely-applicable control strategies. To address this problem, here we develop a new data-driven strategy for the design of shading and natural ventilation controls in residential buildings using policy-based reinforcement learning (RL). To limit undesirable actions and reduce training time, we first used imitation learning to initialize RL training with expert knowledge, yielding an initial policy that reduced simulated late-spring space conditioning loads by≥ 40% in 24 climatically diverse cities. This policy was then trained with RL in four cities representing Mediterranean, semi-arid, humid subtropical, and continental climates. When deployed in cities with unfamiliar yet related climates, these new policies reduced space conditioning loads by≥ 50% in the humid subtropics and by≥ 90% in the other three climates, showing exceptional portability. Further, their performance was unexpectedly robust to variations in dwelling orientation, glazing, internal heat gain, and air leakage. These results show the extraordinary potential of imitation-assisted RL in developing high-performance policies for dynamic passive heating and cooling control that remain effective in unfamiliar situations, removing a substantial barrier to passive systems advancement in carbon-free building operation.