research into Reformulating reinforcement learning
research into Reformulating reinforcement learning
批准号:
2271308
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
未结题
起止时间:
2019 至 --
中文摘要
强化学习(RL)可以被视为在未知环境中的优化,其目标是平衡探索(访问环境中的新状态/动作)和利用(重访具有较大奖励的状态和动作)。近年来,强化学习的力量已经被多次证明,例如AlphaGo击败了世界上最好的围棋选手。本项目的目标是评估最近一种方法[1]对强化学习的适用性,该方法用于表示不确定性的双重性质:随机性和确定性。前者是指通常的概率方法,后者对应于某些东西(例如统计模型中的参数)是固定的但未知的情况。这与RL相关,其中环境可以是真正随机的或简单未知的,例如对手的策略或给定动作的奖励概率。[1]的框架基于外部测量的测量理论概念,自然地弥合了频率论和贝叶斯方法之间的差距。这对于RL来说也是至关重要的,因为这两种方法目前共存。有两个主要的研究问题:1)在性能和计算效率方面,引入更忠实的不确定性表示有什么实际好处?2)为频率论和贝叶斯技术开发的理论保证能扩展到基于外部测量的算法吗?为了回答这些问题,我们首先考虑多武装匪徒的特殊情况。MAB足够简单,可以导出理论保证,同时仍然呈现RL关于探索与利用的基本困境。[1]J·豪西诺arXiv:1801.00569,2018. Naval Group是一家专门从事海军防御的法国公司。MAB和RL解决的顺序资源分配问题出现在军事行动的许多关键方面。例如,军用船只配备了一系列传感器,可以在各种模式下工作;控制这些传感器以实现不同的目标是一个具有挑战性的问题,需要处理不同类型的不确定性出现的复杂环境。研究的背景-强化学习(RL)可以被视为未知环境中的优化,其目标是平衡探索(访问环境中的新状态/行为)和开发(重访国家和行动与大奖励)。研究的目的和目标-这个项目的目的是评估RL的适用性最近的方法的基础上结合的可能性和概率理论。这是相关的RL的环境可以是真正的随机或简单的未知,如与对手的策略或奖励的概率的一个给定的action.The新奇的研究方法-所提出的方法是基于一个新的贝叶斯推理的制定使用的工具的可能性理论模型参数的不确定性。这种方法适用于强化学习问题,其中最初缺乏的知识必须忠实地represented.The潜在的影响,应用程序和好处-通过扩展标准的统计框架,所提出的方法有可能导致新的解决方案MABs和RL一般。在统计学中,考虑的方法可以解释几个标准的统计学,从而提供正式的理由来理解他们的属性和局限性。因此,有可能为现有的技术带来新的见解。研究如何与职权范围有关-拟议的研究是在EPSRC的数学科学主题下,特别是在统计和应用概率以及理论计算机科学下。研究领域:数学科学外部合作伙伴-海军集团
英文摘要
Reinforcement learning (RL) can be viewed as optimisation in an unknown environment, where the goal is to balance exploration (visiting new states/actions in the environment) and exploitation (revisiting states and actions with large rewards). The power of RL has been demonstrated many times in the recent years, for instance with AlphaGo defeating the world's best Go player.The objective of this project is to assess the applicability to RL of a recent approach [1] for representing the dual nature of uncertainty: random and deterministic. The former refers to the usual probabilistic approach and the latter corresponds to the case where something (e.g. a parameter in a statistical model) is fixed but unknown. This is relevant for RL where the environment can be truly random or simply unknown, such as with the strategy of an opponent or with the probability of reward of a given action.The framework of [1], based on the measure-theoretic notion of outer measure, naturally bridges the gap between the frequentist and Bayesian approaches. This can also be crucial for RL where the two approaches currently coexist.There are two main research questions:1)What are the practical benefits of introducing a more faithful representation of uncertainty in terms of performance and computational efficiency?2)Can the theoretical guarantees developed for frequentist and Bayesian techniques be extended to algorithms based on outer measures?To answer these questions, the special case of the multi-armed bandits (MABs) will first be considered. MABs are sufficiently simple to allow for theoretical guarantees to be derived while still presenting the fundamental dilemma of RL regarding exploration vs exploitation.[1] J. Houssineau. Parameter estimation with a class of outer probability measures.arXiv:1801.00569, 2018.Naval Group is a French company specialised in naval-based defence.The sequential resource allocation problems that MABs and RL solve appear in a number of crucial aspects in military operations. For instance, military vessels are equipped with a range of sensors that can operate in various modes; controlling these sensors to fulfil different objectives is a challenging problem that requires dealing with a complex environment where different types of uncertainty arise.The context of the research - Reinforcement learning (RL) can be viewed as optimisation in an unknown environment, where the goal is to balance exploration (visiting new states/actions in the environment) and exploitation (revisiting states and actions with large rewards).The aims and objectives of the research - The objective of this project is to assess the applicability to RL of a recent approach based on a combination of possibility and probability theory. This is relevant for RL where the environment can be truly random or simply unknown, such as with the strategy of an opponent or with the probability of reward of a given action.The novelty of the research methodology - The proposed approach is based on a new formulation of Bayesian inference using the tools of possibility theory to model parameter uncertainty. This approach lends itself to RL problems where the initial absence of knowledge must be faithfully represented.The potential impact, applications, and benefits - By extending the standard statistical framework, the proposed approach has the potential to lead to new solutions for MABs and for RL in general. Within statistics, the considered approach allows for explaining several standard heuristics and hence providing formal ground to understand their properties and limitations. There is therefore a potential for bringing new insights into existing techniques.How the research relates to the remit - The proposed research is under the Mathematical Sciences theme of the EPSRC, in particular under Statistics and Applied Probability and under Theoretical Computer Science.Research Area: Mathematical SciencesExternal Partner - Naval Group
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金