课题基金 / 基金详情

Reinforcement Learning for Finite Horizons (ReLeaF)

Reinforcement Learning for Finite Horizons (ReLeaF)
有限视野强化学习 (ReLeaF)
批准号:
EP/X021513/1
负责人:
Sven Schewe
金额:
$26.0万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

Sven Schewe的其他基金

相似基金

相关文献

中文摘要
翻译
强化学习(RL)是一种学习如何在最初未知的环境中采取行动以优化预期结果的技术,它通过最大化累积奖励的概念来建模。将目标写成时间规范的学习算法有三个关键要素:从规范到适当的有限自动机的转换;将这些有限自动机转化为奖励结构,使提供最优奖励的策略保证提供最优控制;并将其包装成一个折扣方案,对于适当的参数,将确保学习者收敛到最优策略。我们将考虑一种用于自动化和运动规划的流行规范语言的RL问题,即有限视界线性时间时间逻辑ltf。特别是,我们将研究无模型强化学习算法,它比基于模型的算法更适合于难以预测环境行为的现实应用。我们将提出学习算法,提供从有限视界LTL到奖励结构的转换,并对建模为马尔可夫决策过程(mdp)的环境提供满足给定目标的正式保证。我们将把我们的技术扩展到无限状态的mdp,包括可以提供形式保证的变化——比如可数的、有限分支的mdp——并研究我们的技术在更一般的类中提供保证的条件,比如紧凑mdp的平滑保证。我们将通过研究有约束的目标来补充这些研究方向。这是有效地考虑优先目标,其中满足安全约束是优先考虑的,而其他属性(如效率)则被认为是提供相同安全保证的策略之间的决定性因素。
英文摘要
Reinforcement learning (RL) is a technique for learning how to take actions in an initially unknown environment in order to optimise an expected outcome, which is modelled through the notion of maximising an accumulative reward. Learning algorithms with goals written as temporal specifications have three key ingredients: the translation from the specification to appropriate finite automata; the translation of these finite automata to reward structures, such that a strategy that provides optimal rewards is guaranteed to provide optimal control; and a wrapper into a discounting scheme that, for appropriate parameters, will ensure that a learner converge to an optimal strategy.We will consider the RL problems for a popular specification language used in automation and motion planning, the finite horizon linear time temporal logic LTLf. In particular, we will study model-free RL algorithms, which are more suitable to real-world applications where the behaviour of the environment is hard to predict, than its model-based counterpart. We will propose learning algorithms that provide translations from finite horizon LTL to reward structures with formal guarantees of satisfying the given goals for environments modelled as Markov Decision Processes (MDPs). We will extend our techniques to infinite-state MDPs, including variations where formal guarantees can be provided -- like countable, finitely branching MDPs -- and study conditions for our techniques to provide guarantees in more general classes, such as smoothness guarantees for compact MDPs. We will complement these lines of research by looking at goals with constraints. This is effectively considering prioritised goals, where meeting safety constraints takes precedence, while other properties -- such as efficiency -- are considered as tie-breakers among strategies that provide the same safety guarantees.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
Tools and Algorithms for the Construction and Analysis of Systems - 29th International Conference, TACAS 2023, Held as Part of the European Joint Conferences on Theory and Practice of Software, ETAPS 2023, Paris, France, April 22-27, 2023, Proceedings, Part I
系统构建和分析的工具和算法 - 第 29 届国际会议,TACAS 2023,作为欧洲软件理论与实践联合会议的一部分举行,ETAPS 2023,法国巴黎,2023 年 4 月 22-27 日,会议记录,部分
DOI: 10.1007/978-3-031-30823-9_28
发表时间: 2023
期刊:
影响因子: --
作者: [Park S]
通讯作者: Park S
Automated Technology for Verification and Analysis - 21st International Symposium, ATVA 2023, Singapore, October 24-27, 2023, Proceedings, Part I
验证和分析自动化技术 - 第 21 届国际研讨会,ATVA 2023,新加坡,2023 年 10 月 24-27 日,会议记录,第一部分
DOI: 10.1007/978-3-031-45329-8_3
发表时间: 2023
期刊:
影响因子: --
作者: [Li Y]
通讯作者: Li Y
TRUSTED: SecuriTy SummaRies for SecUre SofTwarE Development
  • 批准号:
    EP/X03688X/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $54.33万
  • 财政年份:
    2023
  • 负责人:
    Sven Schewe
  • 依托单位:
Below the Branches of Universal Trees
  • 批准号:
    EP/X017796/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $25.76万
  • 财政年份:
    2023
  • 负责人:
    Sven Schewe
  • 依托单位:
Valuation Structures for Infinite Duration Games
  • 批准号:
    EP/Y027663/1
  • 项目类别:
    Fellowship
  • 资助金额:
    $25.55万
  • 财政年份:
    2023
  • 负责人:
    Sven Schewe
  • 依托单位:
Solving Parity Games in Theory and Practice
  • 批准号:
    EP/P020909/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $52.13万
  • 财政年份:
    2017
  • 负责人:
    Sven Schewe
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: