课题基金 / 基金详情

An Abstraction-based Technique for Safe Reinforcement Learning

An Abstraction-based Technique for Safe Reinforcement Learning
一种基于抽象的安全强化学习技术
批准号:
EP/X015823/1
负责人:
Francesco Belardinelli
金额:
$38.49万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2023
资助国家:
英国
项目状态:
未结题
起止时间:
2023 至 --

项目摘要

项目成果

Francesco Belardinelli的其他基金

相似基金

相关文献

中文摘要
翻译
学习在未知环境中行动的自主代理由于其对人工智能的广泛影响以及其在复杂领域的应用,包括机器人,网络优化和资源分配,一直吸引着研究兴趣。目前,最成功的方法之一是强化学习(RL)。然而,为了学习如何行动,智能体需要探索环境,这在安全关键的场景中意味着它们可能会采取危险的行动,可能会伤害自己,甚至危及人类的生命。因此,在需要同时满足多个安全关键约束的现实应用中,强化学习仍然很少使用。为了缓解这个问题,强化学习算法正在与形式化验证技术相结合,以确保学习的安全性。事实上,形式化方法现在通常应用于复杂系统的规范、设计和验证,因为它们允许获得对其正确和安全行为的类似证明的认证,这意味着对系统工程师和人类用户都是可理解的。这些理想的特征促使采用正式的方法来验证一般的人工智能系统,这些系统被称为安全、可验证、可信赖的人工智能1。尽管如此,将形式化方法应用于人工智能系统会带来重大的新挑战,包括目前使用的大多数机器学习算法的“黑箱”性质。具体到形式化方法在强化学习中的应用,我们确定了当前方法的两个主要缺点,这将在本项目中解决:—大多数当前的验证方法不能很好地扩展应用程序的复杂性。这种状态爆炸问题对于强化学习场景来说尤其严重,在这种场景中,智能体可能不得不在大量的动作/状态转换(例如,自动驾驶汽车)中进行选择。-与单代理设置相比,具有多个学习代理的系统相对较少被探索,因此较少被理解,部分原因是其状态空间的高维性和非平定性。然而,多智能体设置是应用程序的关键,例如自动驾驶汽车的队列行驶和机器人群。为了解决这两个问题,我们提出了一种基于抽象的验证方法,这意味着通过利用系统的对称性来减少状态空间,同时保留其所有与安全相关的特征,从而导致有保证和可扩展的安全行为。该项目设想的研究是及时的,它符合epsrc目前资助的研究组合,因为它符合人工智能和机器人技术的主题,特别是对值得信赖的自主系统的关键战略投资。本提案旨在开发一种可验证安全的强化学习方法,旨在对公众对部署的人工智能解决方案的信任产生积极的社会影响,并促进其在整个社会中的采用。
英文摘要
Autonomous agents learning to act in unknown environments have been attracting research interest due to their wider implications for AI, as well as for their applications in complex domains, including robotics, network optimisation, and resource allocation. Currently, one of the most successful approaches is reinforcement learning (RL). However, to learn how to act, agents are required to explore the environment, which in safety-critical scenarios means that they might take dangerous actions, possibly harming themselves or even putting human lives at risk. Consequently, reinforcement learning is still rarely used in real-world applications, where multiple safety-critical constraints need to be satisfied simultaneously.To alleviate this problem, RL algorithms are being combined with formal verification techniques to ensure safety in learning. Indeed, formal methods are nowadays routinely applied to the specification, design, and verification of complex systems, as they allow to obtain proof-like certification of their correct and safe behaviour, which is meant to be intelligible to system engineers and human users alike. These desirable features have motivated the adoption of formal methods for the verification of general AI systems, which has variously been called safe, verifiable, trustworthy AI 1. Still, the application of formal methods to AI systems raises significant new challenges, including the "black-box" nature of most machine learning algorithms used nowadays. Specific to the application of formal methods to RL, we identify two main shortcomings with current approaches, which will be tackled in this project:- Most of current verification methodologies do not scale well as the complexity of the application increases. This state explosion problem is particularly acute for RL scenarios, where agents might have to chose among a huge number of action/state transitions (e.g., autonomous cars).- Systems with multiple learning agents are comparatively less explored, and therefore less understood, than single-agent settings, partly because of the high-dimensionality of their state-space and their non-stationarity. Yet, multi-agent settings are key for applications, such as platooning for autonomous vehicles and robot swarms.To tackle both problems, we put forward an abstraction-based approach to verification, which is meant to reduce the state space, also by leveraging on symmetries of the system, while preserving all its safety-related features, thus leading to guaranteed and scalable safe behaviours. The research envisaged in this project is timely and it fits with the current portfolio of EPSRC-funded research, as it aligns with the theme of AI and robotics, in particular the key strategic investment in trust-worthy autonomous systems. The present proposal is aimed at developing a verifiably safe RL methodology, which is meant to have a positive societal impact on the trust of the general public towards deployed AI solutions, and to facilitate their adoption within society at large.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Strategy Logics for the Verification of Security Protocols
  • 批准号:
    EP/V009214/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $1.15万
  • 财政年份:
    2021
  • 负责人:
    Francesco Belardinelli
  • 依托单位:
The Third International Workshop on Formal Methods in Artificial Intelligence
  • 批准号:
    EP/V008013/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $0.67万
  • 财政年份:
    2021
  • 负责人:
    Francesco Belardinelli
  • 依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
Incentive and governance schenism study of corporate green washing behavior in China: Based on an integiated view of econfiguration of environmental authority and decoupling logic
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    YU BYUNGJUN
  • 依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
  • 批准号:
    W2433169
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    HAOFEI ZHANG
  • 依托单位:
含Re、Ru先进镍基单晶高温合金中TCP相成核—生长机理的原位动态研究
  • 批准号:
    52301178
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30.00万元
  • 批准年份:
    2023
  • 负责人:
    夏万顺
  • 依托单位: