An Adaptive Robust Dynamic Programming Approach for Decision Making under Model Uncertainty
An Adaptive Robust Dynamic Programming Approach for Decision Making under Model Uncertainty
批准号:
2440945
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
在许多现实世界的问题中,智能体必须在仅部分已知的环境中做出决策。通过与世界的互动,决策者能够获得更多关于系统的信息,从而在未来做出更多有教养的选择。因此,这些问题的一个共同特点是,决策者可以选择之间的决定,导致一个相当无风险,高的即时回报,和更危险的决定,可能会更糟,但可能会提供代理人与以前看不见的信息,他们的环境。在强化学习领域,这种困境通常被称为“探索-利用权衡”(exploration-exploitation trade-off),这是一个活跃的研究领域。理解探索-利用权衡的一个根本挑战是需要测量智能体“学习”的信息增益,并能够理解这些信息如何随着时间的推移而发展。传统上,这可以在贝叶斯框架中完成。然而,贝叶斯框架需要一组初始信念,在实践中,这些信念可能是不精确的。另一种方法是根据最坏情况下的结果做出决策,但是这种方法缺乏考虑学习的能力。在这个项目中,我们的目标是通过考虑一个自适应(即可以包含学习),鲁棒(即在设置中考虑不确定性)的框架来联合收割机两个世界的最佳情况,以解决具有模型不确定性的随机控制问题。我们的出发点是Bielecki等人(2017)的框架,他们考虑了一种自适应的鲁棒方法来解决与投资问题相关的随机控制问题。我们将尝试将他们的方法应用于报童问题。报童问题是一个简单的随机控制问题,涉及学习。在这个问题中,一个代理(报童)必须在观察报纸的销售数量之前选择下一个时期的报纸库存数量,并鼓励了解报纸需求的分布,同时最大限度地减少因未使用的库存或未满足的需求而产生的成本。由于当前的股票选择会影响未来的结果,由于观察到的销售数量的信息的差异,解决这些问题需要了解代理人的信念将如何在未来发生变化。我们希望构造近似参数的最优运输理论的基础上,以减少问题的复杂性。该项目的其他可能目标包括推广目前仅在非常特殊的环境中已知的结果(例如,来自Y. T. Chuang,2019)精确地量化了仅用于学习的库存盈余。对Newsvendor模型的兴趣主要是由于其数学上的易处理性,以及所获得的信息对代理人所做决策的强烈依赖。我们希望这些原则能够更广泛地适用于许多强化学习示例,从而为强化学习的未来发展做出更广泛的贡献。
英文摘要
In many real-world problems an agent must make decisions in an environment that is only partially known. By interacting with the world, the decision-maker is able to obtain more information about the system which allows for more educated choices in the future. Hence, a common characteristic of these problems is that the decision-maker can choose between decisions that lead to a fairly risk-free, high immediate reward, and more risky decisions which may be worse, but may provide the agent with previously unseen information about their environment. In the field of Reinforcement Learning this dilemma is commonly referred to as the "exploration-exploitation trade-off," and is an area of active research.A fundamental challenge in understanding the exploration-exploitation trade-off is that one needs to measure the information gain "learned" by the agent, and to be able to understand how this information develops over time. Classically, this can be done in a Bayesian framework. However, the Bayesian framework requires an initial set of beliefs, and in practice, these may be imprecise. An alternative approach is to make decisions based on outcomes under worst-case scenarios, however this approach lacks the ability to account for learning.In this project we aim to combine the best of both worlds by considering an adaptive (i.e. can incorporate learning), robust (i.e. accounting for uncertainty in the setup) framework for stochastic control problems featuring model uncertainty. Our starting point is the framework of Bielecki et al. (2017), who considered an adaptive, robust approach to a stochastic control problem related to an investment problem. We will attempt to apply their approach to the Newsvendor problem. The Newsvendor problem is a simple stochastic control problem that involves learning. In this problem, an agent (the newsvendor) must choose the number of newspapers to stock for the next period before observing the number of newspapers sold, and is encouraged to learn the distribution of the demand for newspapers, whilst minimising the cost due to unused stock, or unmet demand. As the current choice of stock will affect future outcomes due to differences in information about the number of sales observed, solving such problems requires understanding how the agent's beliefs will change in the future. We hope to construct approximation arguments based on the theory of Optimal Transport in order to reduce the complexity of the problem. Other possible aims of the project include generalising results that are currently known only in very special settings (e.g. from Y.-T. Chuang, 2019) which precisely quantify the surplus in stock used only for the sake of learning.The interest in the Newsvendor model is primarily on account of its mathematical tractability, and the strong dependence of the information acquired on the decisions made by the agent. We expect the principles to be more widely applicable to many RL examples, and may thus contribute more broadly to future developments in Reinforcement Learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
-
批准号:70601028
-
项目类别:青年科学基金项目
-
资助金额:7.0万元
-
批准年份:2006
-
负责人:王明征
-
依托单位:
心理紧张和应力影响下Robust语音识别方法研究
-
批准号:60085001
-
项目类别:专项基金项目
-
资助金额:14.0万元
-
批准年份:2000
-
负责人:韩纪庆
-
依托单位:
ROBUST语音识别方法的研究
-
批准号:69075008
-
项目类别:面上项目
-
资助金额:3.5万元
-
批准年份:1990
-
负责人:高雨青
-
依托单位:
改进型ROBUST序贯检测技术
-
批准号:68671030
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1986
-
负责人:刘有恒
-
依托单位: