Learning to Reason in Reinforcement Learning
Learning to Reason in Reinforcement Learning
批准号:
DP240103278
负责人:
Dr Ehsan Abbasnejad
金额:
$37.67万
依托单位国家:
澳大利亚
项目类别:
Discovery Projects
财政年份:
2024
资助国家:
澳大利亚
项目状态:
未结题
起止时间:
2024-01-01 至 2026-12-31
中文摘要
深度强化学习使用深度神经网络来表示和学习复杂环境中智能主体的最优决策策略。然而,大多数RL方法需要数百万集才能收敛到好的策略,这使得RL很难应用于占用大量资源的现实世界场景。该项目旨在为RL配备反事实推理和结果预测等能力,以显著减少所需的交互次数,改进泛化,并为代理人提供考虑因果关系的能力。这些改进将缩小人工智能和人类能力之间的差距,并扩大RL在现实世界应用中的采用。
英文摘要
Deep Reinforcement Learning (RL) uses deep neural networks to represent and learn optimal decision-making policies for intelligent agents in complex environments. However, most RL approaches require millions of episodes to converge to good policies, making it difficult for RL to be applied in real-world scenarios taking significant resources. This project aims to equip RL with capabilities such as counterfactual reasoning and outcome anticipation to significantly reduce the number of interactions required, improve generalisation, and provide the agent with the capability to consider the cause-effects. These improvements would narrow the gap between AI and human capabilities and broaden the adoption of RL in real-world applications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金