Scalable Inverse Reinforcement Learning Through Multifidelity Bayesian Optimization

Scalable Inverse Reinforcement Learning Through Multifidelity Bayesian Optimization
复制标题

DOI:
10.1109/tnnls.2021.3051012
复制
发表时间:
2021-01-22
影响因子:
10.4
通讯作者:
Ghoreishi, Seyede Fatemeh
Ghoreishi, Seyede Fatemeh
中科院分区:
计算机科学1区
文献类型:
--
作者:
Imani, Mahdi;Ghoreishi, Seyede Fatemeh

文献摘要

被引文献

相似文献

许多实际问题中的数据是根据用户或专家为实现特定目标而做出的决策或行动来获取的。例如,在基因组学和宏基因组学的干预过程中,生物学家头脑中的政策往往反映在这些领域的可用数据中,或者网络物理系统中的数据往往是根据专家/工程师出于控制或稳定等目的而做出的行动/决策而获得的。通过可用数据量化专家的策略,这也被称为奖励函数学习,在逆强化学习(IRL)的背景下已经在文献中广泛讨论。然而,由于以下主要原因,大多数可用的技术都无法处理实际问题:1)缺乏可扩展性:由于现有技术在处理大型系统时的无能或性能差,以及2)缺乏可靠性:由于现有技术在学习过程中无法正确学习最优奖励函数。为此,在这个简短的,我们提出了一个多保真度贝叶斯优化(MMBO)的框架,显着扩展了广泛的现有IRL技术的学习过程。所提出的框架能够合并多个近似器,并有效地考虑到它们的不确定性和计算成本,以平衡学习过程中的探索和利用。通过基因组学,宏基因组学和随机模拟问题集证明了所提出的框架的高性能。
Data in many practical problems are acquired according to decisions or actions made by users or experts to achieve specific goals. For instance, policies in the mind of biologists during the intervention process in genomics and metagenomics are often reflected in available data in these domains, or data in cyber-physical systems are often acquired according to actions/decisions made by experts/engineers for purposes, such as control or stabilization. Quantification of experts' policies through available data, which is also known as reward function learning, has been discussed extensively in the literature in the context of inverse reinforcement learning (IRL). However, most of the available techniques come short to deal with practical problems due to the following main reasons: 1) lack of scalability: arising from incapability or poor performance of existing techniques in dealing with large systems and 2) lack of reliability: coming from the incapability of the existing techniques to properly learn the optimal reward function during the learning process. Toward this, in this brief, we propose a multifidelity Bayesian optimization (MFBO) framework that significantly scales the learning process of a wide range of existing IRL techniques. The proposed framework enables the incorporation of multiple approximators and efficiently takes their uncertainty and computational costs into account to balance exploration and exploitation during the learning process. The proposed framework's high performance is demonstrated through genomics, metagenomics, and sets of random simulated problems.