课题基金 / 基金详情

CAREER: Using Imperfect Predictions to Make Good Decisions

CAREER: Using Imperfect Predictions to Make Good Decisions
职业:利用不完美的预测做出正确的决策
批准号:
1939827
负责人:
Erin Talvitie
金额:
$30.13万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-07-01 至 2023-06-30

项目摘要

项目成果

Erin Talvitie的其他基金

相似基金

相关文献

中文摘要
翻译
当人类和其他动物在世界上航行时,它们在遇到陌生的系统、空间和现象时表现出非凡的灵活性,学会预测自己的行为,并根据这些预测做出正确的决定。这种能力的关键是一个事实,即一个人不需要做出完全准确或完全详细的预测来做出正确的决定。虽然,由于我们天生的局限性,我们对未来的预测必然是有缺陷的,但它们仍然足够有用,足以做出合理的决定。相比之下,对于人工代理人来说,不完美的预测往往会导致决策的灾难性失败。许多现有的方法基本上假设代理人最终会学会做出完美的预测和做出完美的决定,这在足够丰富、复杂的环境中是不合理的。这项工作考虑了开发更能意识到自己的局限性并更健壮的人工代理的问题。能够更有力、更灵活地从真正复杂环境中的经验中学习的代理,有可能影响几乎任何随时间做出决策的应用,例如自主机器人/车辆、个人助理和医疗/法律决策支持。此外,由于该项目将在一所仅限本科生的文科学院进行,本科生研究人员将在这项工作中发挥不可或缺的作用。PI还将加强文科环境的优势,在整个计算机科学课程中加强对关键学科的研究和写作技能的指导。这些技能的明确发展不仅将改善学生为各种职业道路(包括基础研究)做准备,而且还将与扩大该学科参与度的最佳实践保持一致。本项目研究基于模型的强化学习(MBRL),假设代理具有基本的限制,使其无法学习完美的模型或产生最优计划。中心假设是,在这种情况下,MBRL问题不能分解为单独的模型学习和规划问题,每个问题都将对方视为理想化的黑匣子。相反,每个组件的优化过程必须知道它在整个体系结构中的角色,以及它的合作伙伴的限制。这项工作的一个关键目标是推导出与控制性能的真实目标更紧密相关的新的模型质量衡量标准,而不是适应于监督学习设置的一步预测精度的标准衡量标准。另一项是研究如何调整模型学习目标/算法,以解决将使用该模型的特定规划者的限制。此外,还将研究控制算法,通过在基于模型的知识和无模型的知识之间进行调解,有效地利用非均匀质量的模型。最终目标是将这些原则集成到新型MBRL代理中,这些代理对模型类和/或计划器中的限制更加健壮,并且能够在过于复杂和高维的环境中取得成功,这些环境无法准确建模或解决。
英文摘要
As humans and other animals navigate the world they demonstrate remarkable flexibility in encountering unfamiliar systems, spaces and phenomena, learning to make predictions about how they will behave, and making good decisions based on those predictions. Crucial to this ability is the fact that one does not need to make perfectly accurate or fully detailed predictions to make good decisions. Though, due to our natural limitations, our predictions about the future are necessarily flawed, they are nevertheless sufficiently useful to make reasonable decisions. For artificial agents, in contrast, imperfect predictions often lead to catastrophic failures in decision making. Many existing approaches fundamentally assume that the agent will eventually learn to make perfect predictions and make perfect decisions, which is unreasonable in sufficiently rich, complex environments. This work considers the problem of developing artificial agents that are more aware of and more robust to their own limitations. Agents that can more robustly and flexibly learn from experience in truly complex environments have the potential to impact nearly any application in which decisions are made over time, for instance autonomous robots/vehicles, personal assistants, and medical/legal decision support. Furthermore, as the project will be undertaken at an undergraduate-only liberal arts college, undergraduate researchers will play an integral role in the work. The PI will also build on the strength of the liberal arts setting to enhance instruction of key discipline-specific research and writing skills throughout the Computer Science curriculum. Explicit development of these skills will not only improve students' preparation for a wide variety of career paths (including basic research) but is also aligned with best practices for broadening participation in the discipline. This project studies model-based reinforcement learning (MBRL) under the assumption that the agent has fundamental limitations that prevent it from learning a perfect model or from producing optimal plans. The central hypothesis is that in this context the MBRL problem cannot be decomposed into separate model-learning and planning problems, each treating the other as an idealized black box. Rather the optimization process for each component must be aware of its role in the overall architecture and of the limitations of its partner. One key aim of the work is to derive novel measures of model quality that are more tightly related to the true objective of control performance than standard measures of one-step prediction accuracy adapted from supervised learning settings. Another is to investigate how model learning objectives/algorithms can be adapted to account for the limitations of the specific planner that will use the model. Further, control algorithms will be investigated that can make effective use of models of non-homogeneous quality by mediating between model-based and model-free knowledge. The ultimate goal is to integrate these principles into novel MBRL agents that are significantly more robust to limitations in the model class and/or planner and are able to succeed in environments that are too complex and high-dimensional to be modeled or solved exactly.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Using Imperfect Predictions to Make Good Decisions
  • 批准号:
    1552533
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.81万
  • 财政年份:
    2016
  • 负责人:
    Erin Talvitie
  • 依托单位:
国内基金
海外基金
Capture and Release of Droplets Using Advanced Materials for High Technology Applications
  • 批准号:
    52073127
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2020
  • 负责人:
    Alidad Amirfazli
  • 依托单位:
Molecular Interaction Reconstruction of Rheumatoid Arthritis Therapies Using Clinical Data