课题基金 / 基金详情

An Adaptive Robust Dynamic Programming Approach for Decision Making under Model Uncertainty

An Adaptive Robust Dynamic Programming Approach for Decision Making under Model Uncertainty
模型不确定性下决策的自适应鲁棒动态规划方法
批准号:
2440945
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
In many real-world problems an agent must make decisions in an environment that is only partially known. By interacting with the world, the decision-maker is able to obtain more information about the system which allows for more educated choices in the future. Hence, a common characteristic of these problems is that the decision-maker can choose between decisions that lead to a fairly risk-free, high immediate reward, and more risky decisions which may be worse, but may provide the agent with previously unseen information about their environment. In the field of Reinforcement Learning this dilemma is commonly referred to as the "exploration-exploitation trade-off," and is an area of active research.A fundamental challenge in understanding the exploration-exploitation trade-off is that one needs to measure the information gain "learned" by the agent, and to be able to understand how this information develops over time. Classically, this can be done in a Bayesian framework. However, the Bayesian framework requires an initial set of beliefs, and in practice, these may be imprecise. An alternative approach is to make decisions based on outcomes under worst-case scenarios, however this approach lacks the ability to account for learning.In this project we aim to combine the best of both worlds by considering an adaptive (i.e. can incorporate learning), robust (i.e. accounting for uncertainty in the setup) framework for stochastic control problems featuring model uncertainty. Our starting point is the framework of Bielecki et al. (2017), who considered an adaptive, robust approach to a stochastic control problem related to an investment problem. We will attempt to apply their approach to the Newsvendor problem. The Newsvendor problem is a simple stochastic control problem that involves learning. In this problem, an agent (the newsvendor) must choose the number of newspapers to stock for the next period before observing the number of newspapers sold, and is encouraged to learn the distribution of the demand for newspapers, whilst minimising the cost due to unused stock, or unmet demand. As the current choice of stock will affect future outcomes due to differences in information about the number of sales observed, solving such problems requires understanding how the agent's beliefs will change in the future. We hope to construct approximation arguments based on the theory of Optimal Transport in order to reduce the complexity of the problem. Other possible aims of the project include generalising results that are currently known only in very special settings (e.g. from Y.-T. Chuang, 2019) which precisely quantify the surplus in stock used only for the sake of learning.The interest in the Newsvendor model is primarily on account of its mathematical tractability, and the strong dependence of the information acquired on the decisions made by the agent. We expect the principles to be more widely applicable to many RL examples, and may thus contribute more broadly to future developments in Reinforcement Learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
供应链管理中的稳健型(Robust)策略分析和稳健型优化(Robust Optimization )方法研究
  • 批准号:
    70601028
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    7.0万元
  • 批准年份:
    2006
  • 负责人:
    王明征
  • 依托单位:
心理紧张和应力影响下Robust语音识别方法研究
  • 批准号:
    60085001
  • 项目类别:
    专项基金项目
  • 资助金额:
    14.0万元
  • 批准年份:
    2000
  • 负责人:
    韩纪庆
  • 依托单位:
ROBUST语音识别方法的研究
  • 批准号:
    69075008
  • 项目类别:
    面上项目
  • 资助金额:
    3.5万元
  • 批准年份:
    1990
  • 负责人:
    高雨青
  • 依托单位:
改进型ROBUST序贯检测技术
  • 批准号:
    68671030
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1986
  • 负责人:
    刘有恒
  • 依托单位: