Neurosymbolic Transformers for Multi-Agent Communication

Neurosymbolic Transformers for Multi-Agent Communication
复制标题

DOI:
--
复制
发表时间:
2021-01
期刊:
ArXiv
影响因子:
--
通讯作者:
J. Inala;Yichen Yang;James Paulos;Yewen Pu;O. Bastani;Vijay R. Kumar;M. Rinard;Armando Solar-Lezama
J. Inala;Yichen Yang;James Paulos;Yewen Pu;O. Bastani;Vijay R. Kumar;M. Rinard;Armando Solar-Lezama
中科院分区:
其他
文献类型:
--
作者:
J. Inala;Yichen Yang;James Paulos;Yewen Pu;O. Bastani;Vijay R. Kumar;M. Rinard;Armando Solar-Lezama

文献摘要

被引文献

相似文献

我们研究了推断可以解决合作多代理计划问题的沟通结构的问题,同时最大程度地减少了沟通量。我们将通信量量化为通信图的最大程度;该指标捕获了代理商带宽有限的设置。由于决策空间和目标的组合性质,使沟通最小化是具有挑战性的。例如,我们无法通过使用梯度下降训练神经网络来解决此问题。我们提出了一种新型算法,该算法合成了一个控制策略,该策略将用于生成通信图的程序化通信策略与用于选择动作的变压器策略网络相结合。我们的算法首先训练变压器策略,该策略隐含地生成了“软”通信图。然后,它综合了一个“硬化”该图的程序化通信策略,形成了神经肯定变压器。我们的实验表明,我们的方法如何合成产生低级通信图的政策,同时保持近乎最佳的性能。
We study the problem of inferring communication structures that can solve cooperative multi-agent planning problems while minimizing the amount of communication. We quantify the amount of communication as the maximum degree of the communication graph; this metric captures settings where agents have limited bandwidth. Minimizing communication is challenging due to the combinatorial nature of both the decision space and the objective; for instance, we cannot solve this problem by training neural networks using gradient descent. We propose a novel algorithm that synthesizes a control policy that combines a programmatic communication policy used to generate the communication graph with a transformer policy network used to choose actions. Our algorithm first trains the transformer policy, which implicitly generates a"soft"communication graph; then, it synthesizes a programmatic communication policy that"hardens"this graph, forming a neurosymbolic transformer. Our experiments demonstrate how our approach can synthesize policies that generate low-degree communication graphs while maintaining near-optimal performance.