Online planning for multi-agent systems with bounded communication

Online planning for multi-agent systems with bounded communication
复制标题

DOI:
10.1016/j.artint.2010.09.008
复制
发表时间:
2011-02
期刊:
Artif. Intell.
影响因子:
--
通讯作者:
Feng Wu;S. Zilberstein;Xiaoping Chen
Feng Wu;S. Zilberstein;Xiaoping Chen
中科院分区:
其他
文献类型:
--
作者:
Feng Wu;S. Zilberstein;Xiaoping Chen

文献摘要

被引文献

相似文献

提出了一种多智能体环境下不确定环境下的在线规划算法DEC-POMDP。该算法有助于克服离线求解此类问题的高计算复杂性。分散经营的关键挑战是在很少或没有沟通的情况下保持协调的行为,以及在允许沟通的情况下,以最少的沟通来优化价值。该算法通过基于共同知识生成相同的条件计划,并仅在检测到历史不一致时进行通信,从而在必要时允许通信延迟,从而解决了这些挑战。为了适合在线操作,该算法使用一种新的、快速的局部搜索方法来计算良好的局部策略,该方法利用线性规划实现。此外,它限制了每一步使用的内存量,并且可以应用于具有任意水平的问题。实验结果表明,该算法能够解决现有最优离线规划算法无法解决的过大问题,且性能优于最优在线规划方法,在大多数情况下以更少的通信量产生更高的价值。当通信信道不完美(周期性不可用)时,该算法也被证明是有效的。这些结果有助于多智能体环境下决策理论规划的可扩展性。
We propose an online algorithm for planning under uncertainty in multi-agent settings modeled as DEC-POMDPs. The algorithm helps overcome the high computational complexity of solving such problems offline. The key challenges in decentralized operation are to maintain coordinated behavior with little or no communication and, when communication is allowed, to optimize value with minimal communication. The algorithm addresses these challenges by generating identical conditional plans based on common knowledge and communicating only when history inconsistency is detected, allowing communication to be postponed when necessary. To be suitable for online operation, the algorithm computes good local policies using a new and fast local search method implemented using linear programming. Moreover, it bounds the amount of memory used at each step and can be applied to problems with arbitrary horizons. The experimental results confirm that the algorithm can solve problems that are too large for the best existing offline planning algorithms and it outperforms the best online method, producing much higher value with much less communication in most cases. The algorithm also proves to be effective when the communication channel is imperfect (periodically unavailable). These results contribute to the scalability of decision-theoretic planning in multi-agent settings.