Correcting for Selection Bias in Learning-to-rank Systems

Correcting for Selection Bias in Learning-to-rank Systems
复制标题

DOI:
10.1145/3366423.3380255
复制
发表时间:
2020-01
期刊:
Proceedings of The Web Conference 2020
影响因子:
--
通讯作者:
Zohreh Ovaisi;Ragib Ahsan;Yifan Zhang;K. Vasilaky;E. Zheleva
Zohreh Ovaisi;Ragib Ahsan;Yifan Zhang;K. Vasilaky;E. Zheleva
中科院分区:
其他
文献类型:
--
作者:
Zohreh Ovaisi;Ragib Ahsan;Yifan Zhang;K. Vasilaky;E. Zheleva

文献摘要

相似文献

由现代推荐系统收集的点击数据是可用于训练学习排名(LTR)系统的观察数据的重要来源。然而,这些数据受到许多偏差的影响,这些偏差可能导致LTR系统的性能较差。最近用于在这样的系统中进行偏差校正的方法主要集中在位置偏差上,事实上,较高排名的结果(例如,顶级搜索引擎结果)更有可能被点击,即使它们不是给定用户查询的最相关结果。人们对纠正选择偏差的关注较少,这种情况的发生是因为单击的文档反映了最初向用户显示的文档。在这里,我们提出了新的反事实的方法,适应赫克曼的两阶段的方法,并占LTR系统中的选择和位置偏差。我们的经验评估表明,我们提出的方法是更强大的噪声和更好的准确性相比,现有的无偏LTR算法,特别是当有中度或无位置偏差。
Click data collected by modern recommendation systems are an important source of observational data that can be utilized to train learning-to-rank (LTR) systems. However, these data suffer from a number of biases that can result in poor performance for LTR systems. Recent methods for bias correction in such systems mostly focus on position bias, the fact that higher ranked results (e.g., top search engine results) are more likely to be clicked even if they are not the most relevant results given a user’s query. Less attention has been paid to correcting for selection bias, which occurs because clicked documents are reflective of what documents have been shown to the user in the first place. Here, we propose new counterfactual approaches which adapt Heckman’s two-stage method and accounts for selection and position bias in LTR systems. Our empirical evaluation shows that our proposed methods are much more robust to noise and have better accuracy compared to existing unbiased LTR algorithms, especially when there is moderate to no position bias.