课题基金 / 基金详情

Multi-agent Self-improving of Large Language Models (LLMs)

Multi-agent Self-improving of Large Language Models (LLMs)
大型语言模型 (LLM) 的多智能体自我改进
批准号:
2903811
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在快速发展的人工智能(AI)领域,大型语言模型(llm)作为能够理解人类指令并生成有用答案的强大工具脱颖而出。然而,这些模型的发展面临着重大挑战。一般来说,提高法学硕士的生成能力并使其生成与人类价值观保持一致,在很大程度上依赖于大量的人类反馈注释。这种方法虽然有效,但难以扩展,并且可能固有地限制了模型的潜力。作为一种替代方案,一些研究人员转向使用自生成数据(即自学习)来训练法学硕士。自我学习也带来了一系列问题,包括在没有外部纠正的情况下强化现有偏见或不准确的风险。这种困境为在不需要大量人力资源或自学陷阱的情况下提高LLM能力的新方法奠定了基础。本项目试图通过一个多智能体系统提出一个创新的自我完善框架,使这些模型能够通过利用其他同行模型的反馈来学习和增强自己。通过整合各种llm的优势和多样性,该系统有望提高其遵循指令的能力,与人类价值观保持一致,并在最少的人工监督下执行广泛的下游任务。我们的愿景是建立一种可扩展的、有效的方法,通过模型间的交互来持续改进,避开人类反馈的约束和自生成数据训练的局限性。这个自我完善的系统的核心是两个关键问题:1。法学硕士的多样性能否丰富自生成训练数据的质量?2. 不同llm之间的协作能否在确保持续增强的同时减少人工注释的必要性?解决这两个悬而未决的问题可以为人工智能训练/校准方法的新范式打开大门。这一探索旨在促进更有效的人工智能系统开发,减少对人类监督和干预的依赖。因此,这个项目也是对未来人工智能训练策略的开放式探索。它试图通过从严重依赖人类监督的模型转向数据效率更高、自我改进的法学硕士系统,为人工智能社区做出贡献。
英文摘要
In the rapidly evolving field of artificial intelligence (AI), Large Language Models (LLMs) stand out as powerful tools capable of understanding human instructions and generating helpful answers. However, the development of these models faces significant challenges. In general, improving LLMs' generation ability and aligning their generation with human values rely heavily on vast amounts of human feedback annotations. This approach, while effective, is difficult to scale and may inherently limit the models' potential. As an alternative, some researchers turn to train LLMs using self-generated data, i.e., self-learning. Self-learning also presents a set of problems, including the risk of reinforcing existing biases or inaccuracies without external correction. This dilemma sets the stage for a novel approach to advancing LLM capabilities without substantial demand for human resources or the pitfalls of self-learning. This project tries to propose an innovative self-improving framework through a multi-agent system that enables these models to learn and enhance themselves by leveraging feedback from other peer models. By integrating the strengths and diversity of various LLMs, the system is expected to refine its ability to follow instructions, align with human values, and perform across a broad spectrum of downstream tasks with minimal human supervision. The vision is to establish a scalable and efficient method for continuous improvement through inter-model interactions, sidestepping the constraints of human feedback and the limitations of self-generated data training. At the heart of this self-improving system are two pivotal questions: 1. Can the diversity of LLMs enrich the quality of self-generated training data? 2. Can collaboration among different LLMs reduce the necessity for human annotations while ensuring ongoing enhancement? Addressing these two open queries could open the door to a new paradigm in AI training/alignment methodologies. This exploration aims at fostering more efficient AI systems development with reduced reliance on human oversight and intervention. This project, therefore, is also an open-ended exploration into future AI training strategies. It seeks to contribute to the AI community by moving away from heavily human-supervision-dependent models to more data-efficient and self-improving LLM systems.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
基于多模态 AI Agent的面部痤疮瘢痕临床特征评估与治疗方案优化系统的研究
基于首创感染性疾病智能体UNION-Agent的SFTS全流程智慧管理模式探索性研究
  • 批准号:
    JCZRQNB202600735
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
  • 依托单位:
城市随迁老年人活动需求的影响机制与动态空间模型预测研究
  • 批准号:
    2025JJ60227
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    刘瓅珣
  • 依托单位:
基于Agent的自动化渗透测试技术研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    谭劲松
  • 依托单位: