Scalable multi-agent reinforcement learning for distributed control of residential energy flexibility

Scalable multi-agent reinforcement learning for distributed control of residential energy flexibility
复制标题

DOI:
10.1016/j.apenergy.2022.118825
复制
发表时间:
2022-03-22
期刊:
影响因子:
11.2
通讯作者:
McCulloch, Malcolm D.
McCulloch, Malcolm D.
中科院分区:
工程技术1区
文献类型:
--
作者:
Charbonnier, Flora;Morstyn, Thomas;McCulloch, Malcolm D.

文献摘要

被引文献

相似文献

提出了一种新的基于多智能体强化学习的分布式住宅能源协调可扩展类型。合作智能体学习在部分可观测的随机环境中控制电动汽车、空间供暖和灵活负载所提供的灵活性。在标准的独立Q学习方法中,在随机环境中,部分可观测性下的智能体的协调性能会大幅度下降。在这里,从历史数据的离线凸优化中学习和在奖励信号中隔离对总奖励的边际贡献的新颖组合提高了稳定性和规模表现。使用固定大小的Q表,消费者能够评估其对总体系统目标的边际影响,而无需彼此共享个人数据或与中央协调员共享个人数据。案例研究被用来评估探索来源、奖励定义和多主体学习框架的不同组合的适宜性。事实证明,由于减少了能源进口、损失、配电网络拥堵、电池折旧和温室气体排放的成本,拟议的战略在个人和系统层面创造了价值。
This paper proposes a novel scalable type of multi-agent reinforcement learning-based coordination for distributed residential energy. Cooperating agents learn to control the flexibility offered by electric vehicles, space heating and flexible loads in a partially observable stochastic environment. In the standard independent Q-learning approach, the coordination performance of agents under partial observability drops at scale in stochastic environments. Here, the novel combination of learning from off-line convex optimisations on historical data and isolating marginal contributions to total rewards in reward signals increases stability and performance at scale. Using fixed-size Q-tables, prosumers are able to assess their marginal impact on total system objectives without sharing personal data either with each other or with a central coordinator. Case studies are used to assess the fitness of different combinations of exploration sources, reward definitions, and multi-agent learning frameworks. It is demonstrated that the proposed strategies create value at individual and system levels thanks to reductions in the costs of energy imports, losses, distribution network congestion, battery depreciation and greenhouse gas emissions.