LaPraDoR: Unsupervised Pretrained Dense Retriever for Zero-Shot Text Retrieval

LaPraDoR: Unsupervised Pretrained Dense Retriever for Zero-Shot Text Retrieval
复制标题

DOI:
10.48550/arxiv.2203.06169
复制
发表时间:
2022-03
期刊:
--
影响因子:
--
通讯作者:
Canwen Xu;Daya Guo;Nan Duan;Julian McAuley
Canwen Xu;Daya Guo;Nan Duan;Julian McAuley
中科院分区:
其他
文献类型:
--
作者:
Canwen Xu;Daya Guo;Nan Duan;Julian McAuley

文献摘要

相似文献

在本文中,我们提出了一种预训练的双塔密集寻回犬LaPraDoR,它不需要任何监督数据进行训练。具体来说,我们首先提出了迭代对比学习(ICoL),它迭代地训练具有缓存机制的查询和文档编码器。ICoL不仅扩大了负实例的数量,而且还将缓存示例的表示保留在相同的隐藏空间中。然后,我们提出了词典增强密集检索(LEDR)作为一种简单而有效的方法来增强词典匹配的密集检索。我们在最近提出的BEIR基准上对LaPraDoR进行了评估,其中包括9个零射击文本检索任务的18个数据集。实验结果表明,与有监督的密集检索模型相比,LaPraDoR达到了最先进的性能,进一步的分析表明了我们的训练策略和目标的有效性。与重新排序相比,我们的词典增强方法可以在毫秒内运行(快22.5倍),同时实现卓越的性能。
In this paper, we propose LaPraDoR, a pretrained dual-tower dense retriever that does not require any supervised data for training. Specifically, we first present Iterative Contrastive Learning (ICoL) that iteratively trains the query and document encoders with a cache mechanism. ICoL not only enlarges the number of negative instances but also keeps representations of cached examples in the same hidden space. We then propose Lexicon-Enhanced Dense Retrieval (LEDR) as a simple yet effective way to enhance dense retrieval with lexical matching. We evaluate LaPraDoR on the recently proposed BEIR benchmark, including 18 datasets of 9 zero-shot text retrieval tasks. Experimental results show that LaPraDoR achieves state-of-the-art performance compared with supervised dense retrieval models, and further analysis reveals the effectiveness of our training strategy and objectives. Compared to re-ranking, our lexicon-enhanced approach can be run in milliseconds (22.5x faster) while achieving superior performance.