Document Ranking with a Pretrained Sequence-to-Sequence Model

Document Ranking with a Pretrained Sequence-to-Sequence Model
复制标题

DOI:
10.18653/v1/2020.findings-emnlp.63
复制
发表时间:
2020-03
期刊:
--
影响因子:
--
通讯作者:
Rodrigo Nogueira;Zhiying Jiang;Ronak Pradeep;Jimmy J. Lin
Rodrigo Nogueira;Zhiying Jiang;Ronak Pradeep;Jimmy J. Lin
中科院分区:
其他
文献类型:
--
作者:
Rodrigo Nogueira;Zhiying Jiang;Ronak Pradeep;Jimmy J. Lin

文献摘要

被引文献

相似文献

这项工作提出了使用一个预先训练的序列到序列模型的文档排名。我们的方法是从根本上不同于一个普遍采用的基于分类的配方基于编码器的预训练Transformer架构,如BERT。我们展示了如何训练序列到序列模型来生成相关标签作为“目标令牌”,以及如何将这些目标令牌的底层logits解释为用于排名的相关概率。MS MARCO通道排名任务的实验结果表明,我们的排名方法是上级强编码器只模型。在其他三个文档检索测试集合中,我们展示了一种基于零杆传输的方法,该方法优于以前需要域内交叉验证的最先进模型。此外,我们发现,我们的方法显着优于一个编码器,只有架构在数据贫乏的设置。我们通过改变目标令牌来更详细地研究这一观察结果,以探索模型对潜在知识的使用。令人惊讶的是,我们发现目标标记的选择会影响有效性,即使是语义上密切相关的单词。这一发现揭示了为什么我们的序列到序列的文件排名公式是有效的。代码和模型可以在pygaggle.ai上找到。
This work proposes the use of a pretrained sequence-to-sequence model for document ranking. Our approach is fundamentally different from a commonly adopted classification-based formulation based on encoder-only pretrained transformer architectures such as BERT. We show how a sequence-to-sequence model can be trained to generate relevance labels as “target tokens”, and how the underlying logits of these target tokens can be interpreted as relevance probabilities for ranking. Experimental results on the MS MARCO passage ranking task show that our ranking approach is superior to strong encoder-only models. On three other document retrieval test collections, we demonstrate a zero-shot transfer-based approach that outperforms previous state-of-the-art models requiring in-domain cross-validation. Furthermore, we find that our approach significantly outperforms an encoder-only architecture in a data-poor setting. We investigate this observation in more detail by varying target tokens to probe the model’s use of latent knowledge. Surprisingly, we find that the choice of target tokens impacts effectiveness, even for words that are closely related semantically. This finding sheds some light on why our sequence-to-sequence formulation for document ranking is effective. Code and models are available at pygaggle.ai.