课题基金 / 基金详情

Reinforcement learning approach to the optimal stopping problem

Reinforcement learning approach to the optimal stopping problem
最优停止问题的强化学习方法
批准号:
RGPIN-2021-02760
负责人:
Lee, ChiGuhn
金额:
$2.62万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31

项目摘要

项目成果

Lee, ChiGuhn的其他基金

相似基金

相关文献

中文摘要
翻译
我们要解决的是一个研究最多的优化问题,在这个问题中,决策者试图选择一个时间来采取特定的行动,以最大化从随机过程中获得的报酬。这个问题被称为最优停止问题。解决方案应该将给定的决策情况与导致最佳结果的行动对应起来。由于这种映射应该适用于所有可能的情况,因此我们需要找到一种功能作为解决方案,这通常被称为策略。不确定性可能来自多种来源,包括系统状态及其演化、状态变化前的逗留时间、可观测信号中的白噪声等。动作选择中的简单结构-停止与继续-允许在许多情况下最优停止问题的解析解。然而,当状态空间的维度增加时,就像大多数现实情况中经常出现的情况一样,通常会失去最优性,因此必须逐个设计启发式搜索算法。因此,本研究的主要目的是发展最优停止问题的强化学习算法,该算法在高维状态空间和分解框架下是有效的,以便嵌入的最优停止问题可以作为更大问题的一部分来求解。确定了三个具体的目标:(1)将最优停止问题作为一个有监督的学习问题来求解;(2)开发了一个针对最优停止问题的强化学习问题;(2)将一个一般的顺序决策问题分解为一个子问题,从而将最优停止问题作为一个子问题来求解。最优停止问题作为一个独立问题和嵌入问题在从供应链到金融和设备故障检测的广泛领域中无处不在。因此,高效的基于学习的解决方案将以可扩展的方式为从业者提供解决方案。作为一种学习方法,从业者不会充分说明问题的参数。我们解决这个问题的方式确实是创新的。强化学习一直被视为一个具有挑战性的问题,因为优化和估计问题都是混合的。因此,我们将问题视为有监督的学习的方法是创新的,可能会产生影响。
英文摘要
We propose to address one of the most studied optimization problems, in which decision maker tries to choose a time to take a particular action to maximize reward from a stochastic process. This problem is known as the optimal stopping problem. The solution should map a given decision making situation to an action leading to best outcome. As such mapping should be available for all possible situations, we are required to find a function as a solution, which is often called policy. Uncertainties may come from a variety of sources, including system state and its evolution, sojourn times before state change, white noise in the observable signal, and so on. The simple structure in action selection - stop vs. continuation - allows analytical solutions in many instances of the optimal stopping problem. However, when the dimensionality of the state space increases, as often the case in most realistic situations, optimality is usually lost and heuristic search algorithm will have to be designed case by case. Therefore, the main objective of the study is to develop reinforcement learning algorithms for the optimal stopping problem that is efficient with high dimensional state space as well as a decomposition framework so that an embedded optimal stopping problem can be solved as part of a larger problem. Three specific objectives have been identified: (1) solving the optimal stopping problem as a supervised learning problem, (2) developing a reinforcement learning problem that is customized to the optimal stopping problem and (2) decomposing a general sequential decision problem so that an optimal stopping problem can be solved as a sub-problem.  The impact of the proposed problem is likely significant and the proposed approaches are innovative. The ubiquity of the optimal stopping problem as an independent problem and as an embedded problem in a wide range of domains from supply chain to finance and to equipment fault detection. Therefore, efficient learning-based solution will provide solutions to practitioners in a scalable manner. As a learning method the practitioners would not fully specify the parameters of the problem. The way we tackle the problem is truly innovative. Reinforcement learning has been seen as a challenging problem as optimization and estimation problems are all intermingled. Therefore, our approach of seeing the problem as a supervised learning is innovative and likely impactful.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Reinforcement learning approach to the optimal stopping problem
  • 批准号:
    RGPIN-2021-02760
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.62万
  • 财政年份:
    2022
  • 负责人:
    Lee, ChiGuhn
  • 依托单位:
Transfer learning for continual learning in non-stationary environments
  • 批准号:
    553522-2020
  • 项目类别:
    Alliance Grants
  • 资助金额:
    $11.32万
  • 财政年份:
    2021
  • 负责人:
    Lee, ChiGuhn
  • 依托单位:
Machine Learning-enhanced approaches to optimization of supply chain management at Nestlé Canada
  • 批准号:
    538626-2019
  • 项目类别:
    Collaborative Research and Development Grants
  • 资助金额:
    $5.81万
  • 财政年份:
    2020
  • 负责人:
    Lee, ChiGuhn
  • 依托单位:
Transfer learning for continual learning in non-stationary environments
  • 批准号:
    553522-2020
  • 项目类别:
    Alliance Grants
  • 资助金额:
    $11.98万
  • 财政年份:
    2020
  • 负责人:
    Lee, ChiGuhn
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: