InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval

InPars-v2: Large Language Models as Efficient Dataset Generators for Information Retrieval
复制标题

InPars-v2:大型语言模型作为信息检索的高效数据集生成器

DOI:
10.48550/arxiv.2301.01820
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
Rodrigo Nogueira
Rodrigo Nogueira
中科院分区:
--
文献类型:
--
作者:
Vitor Jeronymo;L. Bonifacio;Hugo Abonizio;Marzieh Fadaee;R. Lotufo;Jakub Zavrel;Rodrigo Nogueira

文献摘要

参考文献

被引文献

相似文献

最近,InPars提出了一种在信息检索任务中有效使用大型语言模型(LLM)的方法:通过少量示例,诱导LLM为文档生成相关查询。这些合成的查询-文档对可以用来训练检索器。然而,InPars和最近的Promptagator依赖于专有的llm(如GPT-3和FLAN)来生成此类数据集。在这项工作中,我们引入了InPars-v2,这是一个数据集生成器,它使用开源的llm和现有的强大的重新排序器来选择用于训练的合成查询文档对。一个简单的BM25检索管道,加上一个在InPars-v2数据上进行微调的monoT5重新排序器,可以在BEIR基准测试上获得最新的结果。为了允许研究人员进一步改进我们的方法,我们开放了代码、合成数据和微调模型的源代码:https://github.com/zetaalphavector/inPars/tree/master/tpu
Recently, InPars introduced a method to efficiently use large language models (LLMs) in information retrieval tasks: via few-shot examples, an LLM is induced to generate relevant queries for documents. These synthetic query-document pairs can then be used to train a retriever. However, InPars and, more recently, Promptagator, rely on proprietary LLMs such as GPT-3 and FLAN to generate such datasets. In this work we introduce InPars-v2, a dataset generator that uses open-source LLMs and existing powerful rerankers to select synthetic query-document pairs for training. A simple BM25 retrieval pipeline followed by a monoT5 reranker finetuned on InPars-v2 data achieves new state-of-the-art results on the BEIR benchmark. To allow researchers to further improve our method, we open source the code, synthetic data, and finetuned models: https://github.com/zetaalphavector/inPars/tree/master/tpu
DOI: --
发表时间: 2022-02
期刊: ArXiv
影响因子: --
作者:
L. Bonifacio;Hugo Abonizio;Marzieh Fadaee;Rodrigo Nogueira
通讯作者: L. Bonifacio;Hugo Abonizio;Marzieh Fadaee;Rodrigo Nogueira