Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval

Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval
复制标题

DOI:
10.18653/v1/2022.acl-long.203
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Luyu Gao;Jamie Callan
Luyu Gao;Jamie Callan
中科院分区:
其他
文献类型:
--
作者:
Luyu Gao;Jamie Callan

文献摘要

被引文献

相似文献

最近的研究表明,使用微调的语言模型(LM)的密集检索的有效性。然而,密集的寻回犬很难训练,通常需要经过精心设计的微调管道才能发挥其全部潜力。在本文中,我们确定并解决了密集检索器的两个潜在问题:i)训练数据噪声的脆弱性和ii)需要大批量鲁棒学习嵌入空间。我们使用最近提出的Condenser预训练架构,该架构通过LM预训练学习将信息压缩到密集向量中。在此基础上,我们提出了coCondenser,它增加了一个无监督的语料库级别的对比损失来预热段落嵌入空间。在MS-MARCO、Natural Question和Trivia QA数据集上的实验表明,coCondenser消除了对增强、合成或过滤等繁重数据工程的需求,以及对大批量训练的需求。它显示出与RocketQA相当的性能,RocketQA是一个最先进的,经过大量工程设计的系统,使用简单的小批量微调。
Recent research demonstrates the effectiveness of using fine-tuned language models (LM) for dense retrieval. However, dense retrievers are hard to train, typically requiring heavily engineered fine-tuning pipelines to realize their full potential. In this paper, we identify and address two underlying problems of dense retrievers: i) fragility to training data noise and ii) requiring large batches to robustly learn the embedding space. We use the recently proposed Condenser pre-training architecture, which learns to condense information into the dense vector through LM pre-training. On top of it, we propose coCondenser, which adds an unsupervised corpus-level contrastive loss to warm up the passage embedding space. Experiments on MS-MARCO, Natural Question, and Trivia QA datasets show that coCondenser removes the need for heavy data engineering such as augmentation, synthesis, or filtering, and the need for large batch training. It shows comparable performance to RocketQA, a state-of-the-art, heavily engineered system, using simple small batch fine-tuning.