Reducing Disparate Exposure in Ranking: A Learning To Rank Approach

Reducing Disparate Exposure in Ranking: A Learning To Rank Approach
复制标题

减少排名中的不同曝光:学习排名方法

DOI:
--
复制
发表时间:
2018
期刊:
The Web Conference
影响因子:
--
通讯作者:
Carlos Castillo
Carlos Castillo
中科院分区:
--
文献类型:
--
作者:
Meike Zehlike;Carlos Castillo

文献摘要

被引文献

相似文献

排名搜索结果已经成为我们在线查找内容、产品、地点和人员的主要机制。因此,它们的排序不仅有助于搜索者的满意度,而且还有助于职业和商业机会、教育安置,甚至是被排名者的社会成功。研究人员越来越关注数据驱动的排名模型中的系统性偏差,并提出了各种后处理方法来缓解歧视和机会不平等。然而,这种方法的缺点是,它仍然允许训练不公平的排名模型。在这篇文章中,我们探索了一种新的处理中的方法:DELTR,一个学习到排名的框架,解决了培训时间排名中潜在的歧视和机会不平等的问题。我们根据平均群体曝光率的差异来衡量这些问题,并设计了一个排名器,从相关性和减少这种差异的角度优化搜索结果。我们进行了一项广泛的实验研究,结果表明,从相关性和曝光度的角度来看,“色盲”可能是最好的选择,也可能是最差的选择,这取决于训练集中存在的偏见的程度和种类。我们表明,在所有测试场景中,我们的处理内方法在相关性和曝光率方面都比前处理和后处理方法表现得更好。
Ranked search results have become the main mechanism by which we find content, products, places, and people online. Thus their ordering contributes not only to the satisfaction of the searcher, but also to career and business opportunities, educational placement, and even social success of those being ranked. Researchers have become increasingly concerned with systematic biases in data-driven ranking models, and various post-processing methods have been proposed to mitigate discrimination and inequality of opportunity. This approach, however, has the disadvantage that it still allows an unfair ranking model to be trained. In this paper we explore a new in-processing approach: DELTR, a learning-to-rank framework that addresses potential issues of discrimination and unequal opportunity in rankings at training time. We measure these problems in terms of discrepancies in the average group exposure and design a ranker that optimizes search results in terms of relevance and in terms of reducing such discrepancies. We perform an extensive experimental study showing that being “colorblind” can be among the best or the worst choices from the perspective of relevance and exposure, depending on how much and which kind of bias is present in the training set. We show that our in-processing method performs better in terms of relevance and exposure than a pre-processing and a post-processing method across all tested scenarios.