A Framework for Monte-Carlo Tree Search on CPU-FPGA Heterogeneous Platform via on-chip Dynamic Tree Management

A Framework for Monte-Carlo Tree Search on CPU-FPGA Heterogeneous Platform via on-chip Dynamic Tree Management
复制标题

基于片上动态树管理的 CPU-FPGA 异构平台蒙特卡罗树搜索框架

DOI:
10.1145/3543622.3573177
复制
发表时间:
2023
期刊:
ACM
影响因子:
--
通讯作者:
Prasanna, Viktor
Prasanna, Viktor
中科院分区:
--
文献类型:
--
作者:
Meng, Yuan;Kannan, Rajgopal;Prasanna, Viktor

文献摘要

参考文献

被引文献

相似文献

蒙特卡罗树搜索(MCTS)是人工智能(AI)应用中广泛使用的搜索技术。MCTS管理动态演进的决策树(即,一个其深度和高度在运行时演变的人)来引导AI代理朝向最优策略。入树操作是内存受限的,导致通用处理器上的大规模并行MCTS的关键性能瓶颈。CPU-FPGA加速器可以缓解入树操作的内存瓶颈。然而,现有的FPGA加速器的一个主要挑战是缺乏动态内存管理,由于它们不能有效地支持动态演变的MCTS树。在这项工作中,我们通过提出一个MCTS加速框架来解决这一挑战,该框架(1)结合了算法-硬件协同优化的加速器设计,该加速器设计支持动态演化树上的树内操作,而无需昂贵的硬件重新配置;(2)采用混合并行执行模型,以充分利用CPU-FPGA异构系统中的计算能力;(3)支持基于Python的编程API,以便在运行时将所提出的加速器与RL域特定的基准库轻松集成。我们表明,通过使用我们的框架,我们实现了高达6.8倍的加速比和上级可扩展性的并行工人比国家的最先进的并行MCTS多核系统。
Monte Carlo Tree Search (MCTS) is a widely used search technique in Artificial Intelligence (AI) applications. MCTS manages a dynamically evolving decision tree (i.e., one whose depth and height evolve at run-time) to guide an AI agent toward an optimal policy. In-tree operations are memory-bound leading to a critical performance bottleneck for large-scale parallel MCTS on general-purpose processors. CPU-FPGA accelerators can alleviate the memory bottleneck of in-tree operations. However, a major challenge for existing FPGA accelerators is the lack of dynamic memory management due to which they cannot efficiently support dynamically evolving MCTS trees. In this work, we address this challenge by proposing an MCTS acceleration framework that (1) incorporates an algorithm-hardware co-optimized accelerator design that supports in-tree operations on dynamically evolving trees without expensive hardware reconfiguration; (2) adopts a hybrid parallel execution model to fully exploit the compute power in a CPU-FPGA heterogeneous system; (3) supports Python-based programming API for easy integration of the proposed accelerator with RL domain-specific bench-marking libraries at run-time. We show that by using our framework, we achieve up to 6.8× speedup and superior scalability of parallel workers than state-of-the-art parallel MCTS on multi-core systems.
针对对手的合作问题解决:量化分布式约束满足
DOI: --
发表时间: 2009
期刊:
影响因子: --
作者:
Satomi Baba;Naofumi Nishimura;Atsushi Iwasaki;Makoto Yokoo
通讯作者: Makoto Yokoo
FPGA 上的 Blokus Duo 游戏
DOI: --
发表时间: 2013
期刊: The 17th CSI International Symposium on Computer Architecture & Digital Systems (CADS 2013)
影响因子: --
作者:
A. Jahanshahi;Mohammadkazem Taram;Nariman Eskandari
通讯作者: Nariman Eskandari
并行 MCTS 的管道模式
DOI: --
发表时间: 2018
期刊: International Conference on Agents and Artificial Intelligence
影响因子: --
作者:
S. Mirsoleimani;Jaap van den Herik;A. Plaat;J. Vermaseren
通讯作者: J. Vermaseren
关于UCT的并行化
DOI: --
发表时间: 2007
期刊:
影响因子: --
作者:
T. Cazenave;Nicolas Jouandeau
通讯作者: Nicolas Jouandeau
FPGA 上基于高度可扩展、共享内存、蒙特卡罗树搜索的 Blokus Duo 求解器
DOI: --
发表时间: 2014
期刊: International Conference on Field-Programmable Technology
影响因子: --
作者:
Ehsan Qasemi;Amir Samadi;Mohammad H. Shadmehr;Bardia Azizian;Sajjad Mozaffari;Amir Shirian;B. Alizadeh
通讯作者: B. Alizadeh