课题基金 / 基金详情

CAREER: Structural Estimation and Optimization for Partially Observable Markov Decision Processes and Markov Games

CAREER: Structural Estimation and Optimization for Partially Observable Markov Decision Processes and Markov Games
职业:部分可观察马尔可夫决策过程和马尔可夫博弈的结构估计和优化
批准号:
2236477
负责人:
Yanling Chang
金额:
$52.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-02-01 至 2028-01-31

项目摘要

项目成果

Yanling Chang的其他基金

相似基金

相关文献

中文摘要
翻译
这项学院早期职业发展计划(Career)赠款将通过开发分析方法来加强供应链的安全和风险管理,从而为国家的经济繁荣做出贡献。作为一个特殊的用例,该项目将专注于可持续的海鲜供应链运营。美国是世界上第二大海产品消费国,渔业机构正在寻求对管理做法进行实质性改革,以更好地管理渔业种群动态。管理战略的有效性取决于鱼类种群评估,这种评估受到许多不确定因素和噪声信息来源的影响,非法、未报告和无管制的捕捞活动进一步加剧了这种情况,因为非法、未报告和无管制捕捞活动窃取自然资源,威胁海洋生态系统和海产品供应,并破坏港口和海上安全。该奖项支持开发一种新的框架、分析和算法,可以从不完美的数据中学习渔民和捕鱼对手的偏好和行为,确定调整他们行为的方法,并寻找有效的战略来促进可持续作业和打击IUU捕捞。教育计划将利用类似的方法,从可观察到的数据中得出学生的需求和偏好,并设计有效的教育战略,促进包容性、公平性、多样性和可及性。这项研究将通过与美国海岸警卫队学院和北卡罗来纳州海洋渔业部门的合作来提供信息。此外,包括学生和供应链从业者的数据分析教程和研讨会将提供重要的招聘和推广机会。本研究将探索一个包含结构估计、优化和集成分析的方法论框架,用于不完全信息下的动态决策。目前关于学习和优化动态决策的文献主要假设系统是完全可观察的。虽然已经有大量关于部分可观测马尔可夫决策过程(POMDP)的分析和优化的文献,但本项目将专注于基于可观测历史的POMDP模型基元的逆估计,这是一个研究较少的领域。通过部分可观测马尔可夫博弈(POMG),研究了多方决策过程的优化和逆估计问题。本研究通过发展(I)新的估计方法来从POMDP和POMG的相应数据轨迹中学习POMDP和POMG的模型参数;(Ii)针对具有不精确回报的领导者-追随者POMG的高效求解过程;以及(Iii)基于POMDP和POMG的估计-优化-分析的集成方法,以通过学习、有针对性的调制和适应其他智能体的决策行为来提高代理的性能。这些方法将被应用于改善鱼类种群重建工作,支持国防机构打击IUU捕捞,并确定教育领域课程交付战略的最佳实践。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This Faculty Early Career Development Program (CAREER) grant will contribute to the Nation's economic prosperity by developing analytical methods to enhance security and risk management of supply chains. As a particular use case, the project will focus on sustainable seafood supply chain operations. The US is the second largest consumer of seafood in the world, and fisheries agencies are seeking substantial reforms in management practices to better manage fishery population dynamics. The effectiveness of management strategies hinges upon fish stock assessment which is subject to many sources of uncertainties and noisy information, and is further compounded by illegal, unreported and unregulated (IUU) fishing, which steals natural resources, threatens ocean ecosystems and seafood supply, and undermines port and maritime security. This award supports development of a new framework, analytics, and algorithms that can learn preferences and behavior of fishermen and fishing adversaries from imperfect data, identify ways to modulate their behavior, and search for effective strategies to promote sustainable operations and to combat IUU fishing. The educational plan will utilize similar methods to elicit students' needs and preferences from observable data and to design effective education strategies that promote inclusion, equity, diversity, and accessibility. The research will be informed through collaboration with the US Coast Guard Academy and the North Carolina Division of Marine Fisheries. In addition, tutorials and workshops on data analytics involving both students and supply chain practitioners will provide important recruitment and outreach opportunities.This research will investigate a methodological framework comprising structural estimation, optimization, and integrated analysis for dynamic decision making under imperfect information. The current literature on learning and optimizing dynamic decisions mainly assumes that the system is perfectly observable. While there is an extensive literature on the analysis and optimization for partially observable Markov decision processes (POMDPs), this project will focus on the inverse estimation of the primitives of a POMDP model based upon observable histories, an understudied area. Both optimization and inverse estimation for multi-party decision processes are also considered through partially observable Markov games (POMGs). This research address a knowledge gap by developing (i) new estimation methods to learn model parameters of POMDPs and POMGs from their corresponding data trajectories; (ii) efficient solution procedures for leader-follower POMGs with imprecise reward; and (iii) an integrated methodology of Estimation-Optimization-Analysis based on POMDPs and POMGs to improve an agent's performance by learning, targeted modulating, and adapting to other agents' decision behaviors. The methodologies will be applied to improve fish stock rebuilding efforts, support defense agencies in combating IUU fishing, and identify best practices of course delivery strategies in education.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/tac.2022.3217908
发表时间: 2020-08
期刊: IEEE Transactions on Automatic Control
影响因子: 6.8
作者: [Yanling Chang;Alfredo Garcia;Zhide Wang;Lu Sun]
通讯作者: Yanling Chang;Alfredo Garcia;Zhide Wang;Lu Sun
Dynamic Discrete Choice Estimation with Partially Observable States and Hidden Dynamics
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位: