Not All Relevance Scores are Equal: Efficient Uncertainty and Calibration Modeling for Deep Retrieval Models

Not All Relevance Scores are Equal: Efficient Uncertainty and Calibration Modeling for Deep Retrieval Models
复制标题

DOI:
10.1145/3404835.3462951
复制
发表时间:
2021-05
期刊:
Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Daniel Cohen;Bhaskar Mitra;Oleg Lesota;Navid Rekabsaz;Carsten Eickhoff
Daniel Cohen;Bhaskar Mitra;Oleg Lesota;Navid Rekabsaz;Carsten Eickhoff
中科院分区:
其他
文献类型:
--
作者:
Daniel Cohen;Bhaskar Mitra;Oleg Lesota;Navid Rekabsaz;Carsten Eickhoff

文献摘要

被引文献

相似文献

在任何排名系统中,检索模型都会根据文档与给定搜索查询的相关程度来输出文档的单个分数。虽然检索模型随着日益复杂的架构的引入而不断改进,但很少有工作研究检索模型对超出单个值范围的分数的信念。我们认为,捕获模型相对于其自身对文档评分的不确定性是检索的一个关键方面,它允许在新文档分布、集合中更多地使用当前模型,甚至提高下游任务的有效性。在本文中,我们通过高效的检索模型贝叶斯框架解决了这个问题,该框架通过随机过程捕获模型对相关性得分的信念,同时仅增加可忽略不计的计算开销。我们通过基于排名的校准指标来评估这种信念,表明我们的近似贝叶斯框架通过风险意识重新排名及其置信度校准显着提高了检索模型的排名有效性。最后,我们证明这种额外的不确定性信息对于通过截止预测表示的下游任务是可行且可靠的。
In any ranking system, the retrieval model outputs a single score for a document based on its belief on how relevant it is to a given search query. While retrieval models have continued to improve with the introduction of increasingly complex architectures, few works have investigated a retrieval model's belief in the score beyond the scope of a single value. We argue that capturing the model's uncertainty with respect to its own scoring of a document is a critical aspect of retrieval that allows for greater use of current models across new document distributions, collections, or even improving effectiveness for down-stream tasks. In this paper, we address this problem via an efficient Bayesian framework for retrieval models which captures the model's belief in the relevance score through a stochastic process while adding only negligible computational overhead. We evaluate this belief via a ranking based calibration metric showing that our approximate Bayesian framework significantly improves a retrieval model's ranking effectiveness through a risk aware reranking as well as its confidence calibration. Lastly, we demonstrate that this additional uncertainty information is actionable and reliable on down-stream tasks represented via cutoff prediction.