Reinforcement learning approach to the optimal stopping problem
Reinforcement learning approach to the optimal stopping problem
批准号:
RGPIN-2021-02760
负责人:
Lee, ChiGuhn
金额:
$2.62万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
我们要解决的是一个研究最多的优化问题,在这个问题中,决策者试图选择一个时间来采取特定的行动,以最大化从随机过程中获得的报酬。这个问题被称为最优停止问题。解决方案应该将给定的决策情况与导致最佳结果的行动对应起来。由于这种映射应该适用于所有可能的情况,因此我们需要找到一种功能作为解决方案,这通常被称为策略。不确定性可能来自多种来源,包括系统状态及其演化、状态变化前的逗留时间、可观测信号中的白噪声等。动作选择中的简单结构-停止与继续-允许在许多情况下最优停止问题的解析解。然而,当状态空间的维度增加时,就像大多数现实情况中经常出现的情况一样,通常会失去最优性,因此必须逐个设计启发式搜索算法。因此,本研究的主要目的是发展最优停止问题的强化学习算法,该算法在高维状态空间和分解框架下是有效的,以便嵌入的最优停止问题可以作为更大问题的一部分来求解。确定了三个具体的目标:(1)将最优停止问题作为一个有监督的学习问题来求解;(2)开发了一个针对最优停止问题的强化学习问题;(2)将一个一般的顺序决策问题分解为一个子问题,从而将最优停止问题作为一个子问题来求解。最优停止问题作为一个独立问题和嵌入问题在从供应链到金融和设备故障检测的广泛领域中无处不在。因此,高效的基于学习的解决方案将以可扩展的方式为从业者提供解决方案。作为一种学习方法,从业者不会充分说明问题的参数。我们解决这个问题的方式确实是创新的。强化学习一直被视为一个具有挑战性的问题,因为优化和估计问题都是混合的。因此,我们将问题视为有监督的学习的方法是创新的,可能会产生影响。
英文摘要
We propose to address one of the most studied optimization problems, in which decision maker tries to choose a time to take a particular action to maximize reward from a stochastic process. This problem is known as the optimal stopping problem. The solution should map a given decision making situation to an action leading to best outcome. As such mapping should be available for all possible situations, we are required to find a function as a solution, which is often called policy. Uncertainties may come from a variety of sources, including system state and its evolution, sojourn times before state change, white noise in the observable signal, and so on. The simple structure in action selection - stop vs. continuation - allows analytical solutions in many instances of the optimal stopping problem. However, when the dimensionality of the state space increases, as often the case in most realistic situations, optimality is usually lost and heuristic search algorithm will have to be designed case by case. Therefore, the main objective of the study is to develop reinforcement learning algorithms for the optimal stopping problem that is efficient with high dimensional state space as well as a decomposition framework so that an embedded optimal stopping problem can be solved as part of a larger problem. Three specific objectives have been identified: (1) solving the optimal stopping problem as a supervised learning problem, (2) developing a reinforcement learning problem that is customized to the optimal stopping problem and (2) decomposing a general sequential decision problem so that an optimal stopping problem can be solved as a sub-problem. The impact of the proposed problem is likely significant and the proposed approaches are innovative. The ubiquity of the optimal stopping problem as an independent problem and as an embedded problem in a wide range of domains from supply chain to finance and to equipment fault detection. Therefore, efficient learning-based solution will provide solutions to practitioners in a scalable manner. As a learning method the practitioners would not fully specify the parameters of the problem. The way we tackle the problem is truly innovative. Reinforcement learning has been seen as a challenging problem as optimization and estimation problems are all intermingled. Therefore, our approach of seeing the problem as a supervised learning is innovative and likely impactful.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Reinforcement learning approach to the optimal stopping problem
-
批准号:RGPIN-2021-02760
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.62万
-
财政年份:2022
-
负责人:Lee, ChiGuhn
-
依托单位:
Transfer learning for continual learning in non-stationary environments
-
批准号:553522-2020
-
项目类别:Alliance Grants
-
资助金额:$11.32万
-
财政年份:2021
-
负责人:Lee, ChiGuhn
-
依托单位:
Machine Learning-enhanced approaches to optimization of supply chain management at Nestlé Canada
-
批准号:538626-2019
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$5.81万
-
财政年份:2020
-
负责人:Lee, ChiGuhn
-
依托单位:
Transfer learning for continual learning in non-stationary environments
-
批准号:553522-2020
-
项目类别:Alliance Grants
-
资助金额:$11.98万
-
财政年份:2020
-
负责人:Lee, ChiGuhn
-
依托单位:
Machine Learning-enhanced approaches to optimization of supply chain management at Nestlé Canada
-
批准号:538626-2019
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$4.9万
-
财政年份:2019
-
负责人:Lee, ChiGuhn
-
依托单位:
Assistive sequential decision making framework
-
批准号:RGPIN-2019-05460
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.89万
-
财政年份:2019
-
负责人:Lee, ChiGuhn
-
依托单位:
Data-driven condition-based maintenance models
-
批准号:499283-2016
-
项目类别:Collaborative Research and Development Grants
-
资助金额:$10.77万
-
财政年份:2018
-
负责人:Lee, ChiGuhn
-
依托单位:
Optimal Economic Change Detection with Imperfect Information
-
批准号:RGPIN-2014-04145
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2018
-
负责人:Lee, ChiGuhn
-
依托单位:
Optimal Economic Change Detection with Imperfect Information
-
批准号:RGPIN-2014-04145
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2017
-
负责人:Lee, ChiGuhn
-
依托单位:
Dynamic Optimization with Learning Approach to Dynamic Pricing with Financial Milestones
-
批准号:507238-2016
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2016
-
负责人:Lee, ChiGuhn
-
依托单位:
Optimal Economic Change Detection with Imperfect Information
-
批准号:RGPIN-2014-04145
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2016
-
负责人:Lee, ChiGuhn
-
依托单位:
Optimization of job sequencing into seat production line
-
批准号:484337-2015
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2015
-
负责人:Lee, ChiGuhn
-
依托单位:
Workforce optimization analysis at a retailer
-
批准号:491047-2015
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2015
-
负责人:Lee, ChiGuhn
-
依托单位:
Industrial-academic workshop on optimization in finance and risk
-
批准号:486108-2015
-
项目类别:Regional Office Discretionary Funds
-
资助金额:$0.34万
-
财政年份:2015
-
负责人:Lee, ChiGuhn
-
依托单位:
Optimal Economic Change Detection with Imperfect Information
-
批准号:RGPIN-2014-04145
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2015
-
负责人:Lee, ChiGuhn
-
依托单位:
Optimal Economic Change Detection with Imperfect Information
-
批准号:RGPIN-2014-04145
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.6万
-
财政年份:2014
-
负责人:Lee, ChiGuhn
-
依托单位:
Scheduling system for wind tower manufacturing
-
批准号:473219-2014
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2014
-
负责人:Lee, ChiGuhn
-
依托单位:
Optimal transportation and inventory planning for arctic mining
-
批准号:451446-2013
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2013
-
负责人:Lee, ChiGuhn
-
依托单位:
Optimal product design with flexible BOM
-
批准号:446724-2013
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2013
-
负责人:Lee, ChiGuhn
-
依托单位:
Supply Chain Design and Control in Volatile Economy
-
批准号:249940-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.38万
-
财政年份:2012
-
负责人:Lee, ChiGuhn
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
儿童音乐能力发展对语言与社会认知能力及脑发育的影响
-
批准号:31971003
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:南云
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
多场景网络学习中基于行为-情感-主题联合建模的学习者兴趣挖掘关键技术研究
-
批准号:61702207
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2017
-
负责人:刘智
-
依托单位:
基于异构医学影像数据的深度挖掘技术及中枢神经系统重大疾病的精准预测
-
批准号:61672236
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2016
-
负责人:王骏
-
依托单位: