AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback

AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback
复制标题

DOI:
10.48550/arxiv.2305.14387
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yann Dubois;Xuechen Li;Rohan Taori;Tianyi Zhang;Ishaan Gulrajani;Jimmy Ba;Carlos Guestrin;Percy Liang;Tatsunori Hashimoto
Yann Dubois;Xuechen Li;Rohan Taori;Tianyi Zhang;Ishaan Gulrajani;Jimmy Ba;Carlos Guestrin;Percy Liang;Tatsunori Hashimoto
中科院分区:
其他
文献类型:
--
作者:
Yann Dubois;Xuechen Li;Rohan Taori;Tianyi Zhang;Ishaan Gulrajani;Jimmy Ba;Carlos Guestrin;Percy Liang;Tatsunori Hashimoto

文献摘要

被引文献

相似文献

像ChatGPT这样的大型语言模型(LLM)由于其强大的解释能力而被广泛采用。开发这些LLM涉及复杂但知之甚少的工作流程,需要通过人工反馈进行培训。要复制和理解这种遵循原则的方法,需要解决三个主要挑战:数据收集的高成本、缺乏可信的评估以及缺乏参考方法实现。我们通过AlpacaFarm来应对这些挑战,这是一种模拟器,可以以低成本进行研究和开发,以便从反馈中学习。首先,我们设计LLM提示来模拟人类反馈,比众包工作者便宜50倍,并与人类表现出高度一致。其次,我们提出了一个自动评估,并验证它对人类的指令在现实世界中的互动。第三,我们为从成对反馈中学习的几种方法(PPO,DPO,best-of-n,专家迭代等)提供参考实现。最后,作为对AlpacaFarm的端到端验证,我们在10 k对真实的人类反馈上训练和评估了11个模型,并表明在AlpacaFarm中训练的模型的排名与在人类数据上训练的模型的排名相匹配。作为对AlpacaFarm中可能的研究的展示,我们发现使用奖励模型的方法可以大大改善监督微调,并且我们的参考PPO实现导致对Davinci 003的胜率提高了10%。我们在https://github.com/tatsu-lab/alpaca_farm发布AlpacaFarm的所有组件。
Large language models (LLMs) such as ChatGPT have seen widespread adoption due to their strong instruction-following abilities. Developing these LLMs involves a complex yet poorly understood workflow requiring training with human feedback. Replicating and understanding this instruction-following requires tackling three major challenges: the high cost of data collection, the lack of trustworthy evaluation, and the absence of reference method implementations. We address these challenges with AlpacaFarm, a simulator that enables research and development for learning from feedback at a low cost. First, we design LLM prompts to simulate human feedback that are 50x cheaper than crowdworkers and display high agreement with humans. Second, we propose an automatic evaluation and validate it against human instructions obtained on real-world interactions. Third, we contribute reference implementations for several methods (PPO, DPO, best-of-n, expert iteration, and more) that learn from pairwise feedback. Finally, as an end-to-end validation of AlpacaFarm, we train and evaluate eleven models on 10k pairs of real human feedback and show that rankings of models trained in AlpacaFarm match rankings of models trained on human data. As a demonstration of the research possible in AlpacaFarm, we find that methods that use a reward model can substantially improve over supervised fine-tuning and that our reference PPO implementation leads to a +10% improvement in win-rate against Davinci003. We release all components of AlpacaFarm at https://github.com/tatsu-lab/alpaca_farm.