Branes with brains: exploring string vacua with deep reinforcement learning

Branes with brains: exploring string vacua with deep reinforcement learning
复制标题

DOI:
10.1007/jhep06(2019)003
复制
发表时间:
2019-03
影响因子:
5.4
通讯作者:
James Halverson;B. Nelson;Fabian Ruehle
James Halverson;B. Nelson;Fabian Ruehle
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
James Halverson;B. Nelson;Fabian Ruehle

文献摘要

被引文献

相似文献

我们提出深度强化学习作为探索弦真空景观的无模型方法。作为一个具体的应用,我们利用称为异步优势演员批评家的人工智能代理来探索具有相交 D6 膜的 IIA 型紧化。当通过改变 D6 膜配置来探索不同的弦背景配置时,代理会收到与弦一致性条件和与标准模型真空的接近度相关的奖励和惩罚。这些反过来又被用来更新代理的策略和价值神经网络,以改善其行为。通过强化学习,智能体在这两项任务中的表现都得到了显着提高,并且对于某些任务,它比随机游走器找到了更多解决方案。在一种情况下,我们证明代理学习了一种人类衍生的策略来寻找一致的字符串模型。在另一种情况下,不存在人类衍生的策略,智能体会学习一种真正的新策略,该策略在单位时间内以两倍的效率实现相同的目标。我们的结果表明,智能体学习同时解决各种弦理论一致性条件,这些条件用非线性耦合丢番图方程来表达。
We propose deep reinforcement learning as a model-free method for exploring the landscape of string vacua. As a concrete application, we utilize an artificial intelligence agent known as an asynchronous advantage actor-critic to explore type IIA compactifications with intersecting D6-branes. As different string background configurations are explored by changing D6-brane configurations, the agent receives rewards and punishments related to string consistency conditions and proximity to Standard Model vacua. These are in turn utilized to update the agent’s policy and value neural networks to improve its behavior. By reinforcement learning, the agent’s performance in both tasks is significantly improved, and for some tasks it finds a factor ofmore solutions than a random walker. In one case, we demonstrate that the agent learns a human-derived strategy for finding consistent string models. In another case, where no human-derived strategy exists, the agent learns a genuinely new strategy that achieves the same goal twice as efficiently per unit time. Our results demonstrate that the agent learns to solve various string theory consistency conditions simultaneously, which are phrased in terms of non-linear, coupled Diophantine equations.