Query-driven Segment Selection for Ranking Long Documents

Query-driven Segment Selection for Ranking Long Documents
复制标题

DOI:
10.1145/3459637.3482101
复制
发表时间:
2021-09
期刊:
Proceedings of the 30th ACM International Conference on Information & Knowledge Management
影响因子:
--
通讯作者:
Youngwoo Kim;Razieh Rahimi;Hamed Bonab;James Allan
Youngwoo Kim;Razieh Rahimi;Hamed Bonab;James Allan
中科院分区:
其他
文献类型:
--
作者:
Youngwoo Kim;Razieh Rahimi;Hamed Bonab;James Allan

文献摘要

相似文献

基于transformer的排名显示出最先进的性能。然而,他们的自我注意操作大多无法处理长序列。训练这些排名器的常见方法之一是选择每个文档的一些片段作为训练数据,例如第一个片段。但是,这些段可能不包含文档中与查询相关的部分。为了解决这个问题,我们提出了从长文档中选择查询驱动的片段来构建训练数据。段选择器为相关样本提供更准确的标签,为不相关样本提供更难预测的标签。实验结果表明,与建议的段选择器训练的基本BERT的排序器显着优于训练的启发式选择的段,并表现出等同于最先进的模型与本地化的自我注意,可以处理更长的输入序列。我们的研究结果开辟了新的方向,设计高效的变压器为基础的排名。
Transformer-based rankers have shown state-of-the-art performance. However, their self-attention operation is mostly unable to process long sequences. One of the common approaches to train these rankers is to heuristically select some segments of each document, such as the first segment, as training data. However, these segments may not contain the query-related parts of documents. To address this problem, we propose query-driven segment selection from long documents to build training data. The segment selector provides relevant samples with more accurate labels and non-relevant samples which are harder to be predicted. The experimental results show that the basic BERT-based ranker trained with the proposed segment selector significantly outperforms that trained by the heuristically selected segments, and performs equally to the state-of-the-art model with localized self-attention that can process longer input sequences. Our findings open up new direction to design efficient transformer-based rankers.