Learning, Innovation, and Explanation in Self-Organising Multi-Agent Systems
Learning, Innovation, and Explanation in Self-Organising Multi-Agent Systems
批准号:
2127915
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
最初的研究方案侧重于多智能体系统中的学习、创新和解释。关于多智能体系统的设计,已经确定了三个层次:智能体(个体)层、智能体间(社会)层和系统层。以前的工作已经提出了受社会启发的机制,以促进主体间水平的集体行动和自我组织,以及基于GP的技术,用于在主体和系统水平上的适应和近似优化。我们希望探索这三个层次上的学习和创新,以此作为应对环境潜在动态变化和确保可持续发展的方法。我们注意到,可以应用于主体级别的方法,例如用于策略优化的强化学习,与如果我们将系统实体视为控制主体,可以应用于系统层次的方法是相同的。到目前为止,我们的研究重点是构建一个单一的系统范围的策略,所有代理都必须遵守。然而,从多中心的角度来看,代理人应该能够制定自己的政策和行动计划,并参与制定系统政策。这将需要一种机制来结合不同的感官经验和观点,以集体学习的方式使用通信和知识转移协议,以实现社会学习和社会创新。我们在这个阶段提出的研究问题是:应该如何制定和解释个人和集体系统政策,以便即使在动态环境中也能可持续地改善和维持系统性能和代理满意度?公式指的是分布式学习算法和协议,使代理能够构造单独的和集体的策略,以及策略的表示方式。可解释性指的是导致创建新策略的学习算法和策略本身的可理解性。对于系统设计人员和用户来说,了解为什么选择了一个策略,以及为什么它比也可以选择的其他候选策略更好,从而能够信任它的建议并收集关于问题领域的有价值的知识是很重要的。有些学习模型本质上是可理解的,而另一些则不是,有些不可理解的模型可能会变得可理解、可解释和可解释。可理解的模型可以为人类本质上难以理解的问题领域提供新的曙光,并导致采用新的更好的策略来搜索建立策略所依据的规则集的理论无限空间。在这一阶段,我们确定了以下高层次的任务:-回顾以前关于多代理系统设计的工作,重点放在上面确定的三个层面:代理、社会和系统层面--确定以前工作中的知识差距,即关于政策的多中心构建和可解释性--根据文献综述和关于知识差距的结论撰写调查报告--确定可能的应用程序和系统,作为政策构建和选择的运行范例--反复地为三个层面的政策构建提出解决方案:单一系统政策(当前工作);多代理政策;多智能体策略和集体制定的系统策略本项目旨在为集体自适应系统的设计提出一种新的范例,其中系统策略不是手工制作的,设计者不是硬编码不变的规则集,而是自动构造的结果,明确的目标是提高系统性能和智能体满意度。EPSRC的研究领域:控制工程、工程设计、人工智能技术、软件工程
英文摘要
The initial research proposal focuses on learning, innovation, and explanation in multi-agent systems. Three levels have been identified concerning the design of multi-agent systems: the agent (individual) level, the inter-agent (social) level, and the system level. Previous work has proposed socially-inspired mechanisms to promote collective action and self-organisation at the inter-agent level and GP-based techniques for adaptation and approximate optimisation at the agent and system levels. We would like to explore learning and innovation at all three levels as ways to handle potentially dynamic changes in the environment and ensure sustainability.We note that the methods which can be applied to the agent level, such as reinforcement learning for policy optimisation, are the same sort of methods which can be applied to the system level if we consider the system entity as a control agent. So far, our research has focused on constructing a single system-wide policy by which all agents must abide. However, in a polycentric perspective, the agents should be able to formulate their own policies and plans of action and to participate in the formulation of system policies. This would require mechanisms for combining different sensory experiences and opinions, using communication and knowledge transfer protocols in a collective learning fashion in order to achieve social learning and social innovation.The research question we propose at this stage is: how should individual and collective system policies be formulated and explained so as to sustainably improve and maintain system performance and agent satisfaction, even in dynamic environments? Formulation refers both to the distributed learning algorithms and protocols which enable the agents to construct individual and collective policies and to the way the policies are represented. Explainability refers to the intelligibility of both the learning algorithms which result in the creation of new policies and the policies themselves. Understanding why a policy has been selected and why it is better than other candidate policies which could also have been selected is important for system designers and users to be able to trust its recommendations and gather valuable knowledge regarding the problem domain. Some learning models are inherently intelligible while others are not and it is possible that some unintelligible models can be made understandable, explainable, and interpretable. Intelligible models could shed a new light upon problem domains which are intrinsically hard for humans to understand and lead to the adoption of new and better strategies for searching the theoretically infinite space of sets of rules upon which policies are built. At this stage we have identified the following high-level tasks:- Review previous work on the design of multi-agent systems, focusing on the three levels identified above: the agent, the social, and the system levels- Identify knowledge gaps in previous work, namely regarding polycentric construction of policies and explainability- Write survey paper based on both the literature review and the conclusions about knowledge gaps- Identify possible applications and systems to use as running examples for policy construction and selection- Iteratively propose solutions to three levels of policy construction: single system policy (current work); multiple agent policies; multiple agent policies and a collectively formulated system policyThis project aims at proposing a new paradigm for the design of collective adaptive systems in which system policies are not handcrafted and designers do not hardcode immutable rule sets, but rather the result of automatic construction with the explicit goal of improving system performance and agent satisfaction.EPSRC Research Areas: Control Engineering, Engineering design, Artificial intelligence technologies, Software engineering
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金