Scalable Exploration for Neural Online Learning to Rank with Perturbed Feedback

Scalable Exploration for Neural Online Learning to Rank with Perturbed Feedback
复制标题

DOI:
10.1145/3477495.3532057
复制
发表时间:
2022-06
期刊:
Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Yiling Jia;Hongning Wang
Yiling Jia;Hongning Wang
中科院分区:
其他
文献类型:
--
作者:
Yiling Jia;Hongning Wang

文献摘要

相似文献

深度神经网络(DNN)在提高检索任务的排名性能方面表现出显著的优势。在DNN优化和泛化的最新发展的推动下,通过与用户的交互在线学习神经排名模型成为可能。然而,模型学习所需的探索必须在整个神经网络参数空间中进行,这是非常昂贵的,并限制了这种在线解决方案在实践中的应用。在这项工作中,我们提出了一个有效的探索策略,在线交互式神经排名学习的基础上自举。我们的解决方案是基于扰动用户点击反馈训练的排名模型的集合。所提出的方法消除了显式置信集的构造和相关的计算开销,这使得在线神经排序器训练能够在理论保证的情况下在实践中有效地执行。在两个公共学习排名基准数据集上与一系列最先进的OL2R算法进行了广泛的比较,证明了我们提出的神经OL2R解决方案的有效性和计算效率。
Deep neural networks (DNNs) demonstrates significant advantages in improving ranking performance in retrieval tasks. Driven by the recent developments in optimization and generalization of DNNs, learning a neural ranking model online from its interactions with users becomes possible. However, the required exploration for model learning has to be performed in the entire neural network parameter space, which is prohibitively expensive and limits the application of such online solutions in practice. In this work, we propose an efficient exploration strategy for online interactive neural ranker learning based on bootstrapping. Our solution is based on an ensemble of ranking models trained with perturbed user click feedback. The proposed method eliminates explicit confidence set construction and the associated computational overhead, which enables the online neural rankers training to be efficiently executed in practice with theoretical guarantees. Extensive comparisons with an array of state-of-the-art OL2R algorithms on two public learning to rank benchmark datasets demonstrate the effectiveness and computational efficiency of our proposed neural OL2R solution.