research into Reformulating reinforcement learning
research into Reformulating reinforcement learning
批准号:
2271308
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
未结题
起止时间:
2019 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Reinforcement learning (RL) can be viewed as optimisation in an unknown environment, where the goal is to balance exploration (visiting new states/actions in the environment) and exploitation (revisiting states and actions with large rewards). The power of RL has been demonstrated many times in the recent years, for instance with AlphaGo defeating the world's best Go player.The objective of this project is to assess the applicability to RL of a recent approach [1] for representing the dual nature of uncertainty: random and deterministic. The former refers to the usual probabilistic approach and the latter corresponds to the case where something (e.g. a parameter in a statistical model) is fixed but unknown. This is relevant for RL where the environment can be truly random or simply unknown, such as with the strategy of an opponent or with the probability of reward of a given action.The framework of [1], based on the measure-theoretic notion of outer measure, naturally bridges the gap between the frequentist and Bayesian approaches. This can also be crucial for RL where the two approaches currently coexist.There are two main research questions:1)What are the practical benefits of introducing a more faithful representation of uncertainty in terms of performance and computational efficiency?2)Can the theoretical guarantees developed for frequentist and Bayesian techniques be extended to algorithms based on outer measures?To answer these questions, the special case of the multi-armed bandits (MABs) will first be considered. MABs are sufficiently simple to allow for theoretical guarantees to be derived while still presenting the fundamental dilemma of RL regarding exploration vs exploitation.[1] J. Houssineau. Parameter estimation with a class of outer probability measures.arXiv:1801.00569, 2018.Naval Group is a French company specialised in naval-based defence.The sequential resource allocation problems that MABs and RL solve appear in a number of crucial aspects in military operations. For instance, military vessels are equipped with a range of sensors that can operate in various modes; controlling these sensors to fulfil different objectives is a challenging problem that requires dealing with a complex environment where different types of uncertainty arise.The context of the research - Reinforcement learning (RL) can be viewed as optimisation in an unknown environment, where the goal is to balance exploration (visiting new states/actions in the environment) and exploitation (revisiting states and actions with large rewards).The aims and objectives of the research - The objective of this project is to assess the applicability to RL of a recent approach based on a combination of possibility and probability theory. This is relevant for RL where the environment can be truly random or simply unknown, such as with the strategy of an opponent or with the probability of reward of a given action.The novelty of the research methodology - The proposed approach is based on a new formulation of Bayesian inference using the tools of possibility theory to model parameter uncertainty. This approach lends itself to RL problems where the initial absence of knowledge must be faithfully represented.The potential impact, applications, and benefits - By extending the standard statistical framework, the proposed approach has the potential to lead to new solutions for MABs and for RL in general. Within statistics, the considered approach allows for explaining several standard heuristics and hence providing formal ground to understand their properties and limitations. There is therefore a potential for bringing new insights into existing techniques.How the research relates to the remit - The proposed research is under the Mathematical Sciences theme of the EPSRC, in particular under Statistics and Applied Probability and under Theoretical Computer Science.Research Area: Mathematical SciencesExternal Partner - Naval Group
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金