InPars: Data Augmentation for Information Retrieval using Large Language Models

InPars: Data Augmentation for Information Retrieval using Large Language Models
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
L. Bonifacio;Hugo Abonizio;Marzieh Fadaee;Rodrigo Nogueira
L. Bonifacio;Hugo Abonizio;Marzieh Fadaee;Rodrigo Nogueira
中科院分区:
其他
文献类型:
--
作者:
L. Bonifacio;Hugo Abonizio;Marzieh Fadaee;Rodrigo Nogueira

文献摘要

被引文献

相似文献

信息检索社区最近见证了一场革命,由于大型预训练的Transformer模型。这场革命的另一个关键因素是MS MARCO数据集,其规模和多样性使零触发迁移学习能够应用于各种任务。然而,并不是所有的IR任务和领域都可以从一个单一的数据集中平等地受益。对各种NLP任务的广泛研究表明,使用特定领域的训练数据,而不是通用数据,可以提高神经模型的性能。在这项工作中,我们利用大型预训练语言模型的少量功能作为IR任务的合成数据生成器。我们表明,仅在无监督数据集上进行微调的模型优于BM 25等强基线以及最近提出的自监督密集检索方法。此外,在监督数据和我们的合成数据上进行微调的检索器比仅在监督数据上进行微调的模型实现了更好的零射击转移。代码、模型和数据可在https://github.com/zetaalphavector/inpars上获得。
The information retrieval community has recently witnessed a revolution due to large pretrained transformer models. Another key ingredient for this revolution was the MS MARCO dataset, whose scale and diversity has enabled zero-shot transfer learning to various tasks. However, not all IR tasks and domains can benefit from one single dataset equally. Extensive research in various NLP tasks has shown that using domain-specific training data, as opposed to a general-purpose one, improves the performance of neural models. In this work, we harness the few-shot capabilities of large pretrained language models as synthetic data generators for IR tasks. We show that models finetuned solely on our unsupervised dataset outperform strong baselines such as BM25 as well as recently proposed self-supervised dense retrieval methods. Furthermore, retrievers finetuned on both supervised and our synthetic data achieve better zero-shot transfer than models finetuned only on supervised data. Code, models, and data are available at https://github.com/zetaalphavector/inpars .