课题基金 / 基金详情

Learning and Search in Decision Domains Featuring Large Action Sets and Uncertainty

Learning and Search in Decision Domains Featuring Large Action Sets and Uncertainty
具有大型动作集和不确定性的决策域中的学习和搜索
批准号:
RGPIN-2018-06677
负责人:
Buro, Michael
金额:
$2.99万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2019
资助国家:
加拿大
项目状态:
已结题
起止时间:
2019-01-01 至 2020-12-31

项目摘要

项目成果

Buro, Michael的其他基金

相似基金

相关文献

中文摘要
翻译
人工智能(AI)研究已经取得了长足的进步,创造出了在决策领域挑战人类霸权的系统,比如国际象棋、危险边缘、股票交易,以及最近的图像识别、雅达利2600街机游戏和亚洲棋盘游戏围棋。相比之下,在流行的电子游戏中,AI的发展往往是缓慢的,这些游戏通常具有大的动作空间、实时限制、多玩家和隐藏信息,在许多情况下,人类专家仍然可以轻松地超越最好的AI系统。***人类在这些领域的优势可以部分归因于我们在保持解决方案的同时简化问题的能力,在不同的抽象层次上进行搜索的能力(例如,仅在高级解决方案概念似乎不起作用时才查看细节),从观察到的行为推断意图,以及快速调整对手和合作伙伴的能力。上面列出的方法有助于创建强大的AI系统。例如,训练策略网络和使用蒙特卡罗搜索来确定良好的低级动作,目前还不足以在具有大型动作空间和由具有微观效果的动作组成的长时间播放情节的领域中达到人类专家水平的表现。为了克服这些问题,我们建议研究如何更好地将启发式搜索(可以通过向前看来评估行动的优点)与机器学习结合起来,以处理大型组合行动空间、不确定性和代理合作。主要的长期研究目标是:1)使用深度神经网络从自我博弈中学习分层策略;2)理解启发式搜索与学习策略在前向模型不可用的领域中的作用;3)合作多智能体领域的学习策略;4)合作和开发的数据高效智能体建模。为了实现这些长期目标,我们从简单的任务开始,包括从人类训练数据中进行监督学习,研究中等行动空间领域的强化学习,以及将人类合作策略集成到现有的基于搜索的人工智能系统中。***在本提案的目标领域取得实质性进展将对技术和社会产生深远影响。在一个机器可以学会在多智能体环境中表现良好,并能制定和执行有效的高级行动计划的世界里,我们可能离一般的类人智能只有一步之遥。
英文摘要
Artificial Intelligence (AI) research has come a long way creating systems that challenge human supremacy in decision domains such as Chess, Jeopardy, stock trading, and recently image recognition, Atari 2600 arcade games, and the Asian boardgame Go. By contrast, AI progress in popular video games which often feature large action spaces, real-time constraints, multiple players, and hidden information, has been slow, and in many cases human experts can still easily outperform the best AI systems.***The human advantage in these domains can in part be attributed to our abilities to simplify problems while maintaining solutions, to search at different abstraction levels (e.g., looking into details only when high-level solution concepts do not seem to work), to infer intentions from observed actions, and to quickly adjust to opponents and partners. The methods that have been instrumental to creating strong AI systems listed above. For example, training policy networks and using Monte Carlo search to determine good low-level actions are currently not powerful enough to achieve human expert level performance in domains featuring large action spaces and long playing episodes consisting of actions with microscopic effects.***To overcome these problems, we propose to investigate how to better integrate heuristic search (which can evaluate the merit of actions by looking ahead) with machine learning to deal with large combinatorial action spaces, uncertainty, and agent cooperation. The main long-term research objectives are: 1) learning hierarchical policies from self-play using deep neural networks, 2) understanding the role of heuristic search vs. learned policies in domains for which forward-models are not available, 3) learning strategies in cooperative multi-agent domains, and 4) data efficient agent modelling for cooperation and exploitation. We approach these long-term goals by starting with simpler tasks involving supervised learning from human training data, studying reinforcement learning in medium-sized action space domains, and integrating human cooperation strategies into existing search-based AI systems.***Making substantial progress in the target domains of this proposal will have a profound impact on technology and society. In a world in which machines can learn to perform well in multi-agent settings and can formulate and execute effective high-level action plans, we may just be a step away from general human-like intelligence.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Learning and Search in Decision Domains Featuring Large Action Sets and Uncertainty
  • 批准号:
    RGPIN-2018-06677
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $5.97万
  • 财政年份:
    2022
  • 负责人:
    Buro, Michael
  • 依托单位:
Learning and Search in Decision Domains Featuring Large Action Sets and Uncertainty
  • 批准号:
    RGPIN-2018-06677
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2021
  • 负责人:
    Buro, Michael
  • 依托单位:
Learning and Search in Decision Domains Featuring Large Action Sets and Uncertainty
  • 批准号:
    RGPIN-2018-06677
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2020
  • 负责人:
    Buro, Michael
  • 依托单位:
Learning and Search in Decision Domains Featuring Large Action Sets and Uncertainty
  • 批准号:
    RGPIN-2018-06677
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2018
  • 负责人:
    Buro, Michael
  • 依托单位:
海外基金