Monte Carlo Estimates of Evaluation Metric Error and Bias

Monte Carlo Estimates of Evaluation Metric Error and Bias
复制标题

DOI:
10.18122/cs_facpubs/148/boisestate
复制
发表时间:
2018-08
期刊:
Computer Science Faculty Publications and Presentations
影响因子:
--
通讯作者:
Mucun Tian;Michael D. Ekstrand
Mucun Tian;Michael D. Ekstrand
中科院分区:
其他
文献类型:
--
作者:
Mucun Tian;Michael D. Ekstrand

文献摘要

相似文献

传统的离线推荐系统评估应用机器学习和信息检索的指标,而这些指标的基本假设不再成立。这导致在衡量top-N推荐性能(如精度、召回率和nDCG)时出现明显的误差和偏差。这些错误的几个具体原因,包括流行偏差和错误分类的诱饵项目,在现有文献中得到了很好的探讨。在本文中,我们调查了一系列识别和解决这些问题的工作,并报告了我们正在进行的工作,以模拟推荐数据生成和评估过程,以量化评估度量误差的程度,并评估其对各种假设的敏感性
Traditional offline evaluations of recommender systems apply metrics from machine learning and information retrieval in settings where their underlying assumptions no longer hold. This results in significant error and bias in measures of top-N recommendation performance, such as precision, recall, and nDCG. Several of the specific causes of these errors, including popularity bias and mis-classified decoy items, are well-explored in the existing literature. In this paper we survey a range of work on identifying and addressing these problems, and report on our work in progress to simulate the recommender data generation and evaluation processes to quantify the extent of evaluation metric errors and assess their sensitivity to various assumptions