Inference at Scale: Significance Testing for Large Search and Recommendation Experiments

Inference at Scale: Significance Testing for Large Search and Recommendation Experiments
复制标题

DOI:
10.1145/3539618.3592004
复制
发表时间:
2023-05
期刊:
Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Ngozi Ihemelandu;Michael D. Ekstrand
Ngozi Ihemelandu;Michael D. Ekstrand
中科院分区:
其他
文献类型:
--
作者:
Ngozi Ihemelandu;Michael D. Ekstrand

文献摘要

相似文献

已经进行了一些信息检索研究,以评估哪些统计技术适用于比较系统。然而,这些研究都集中在TREC式的实验,通常只有不到100个主题。对于大型搜索和推荐实验,没有类似的工作;这类研究通常有数千个主题或用户,相关性判断要稀疏得多,因此不清楚分析传统TREC实验的建议是否适用于这些设置。在本文中,我们实证研究的行为显着性测试与大搜索和推荐评价数据。我们的研究结果表明,Wilcoxon和Sign检验在大样本量下的1型错误率显著高于bootstrap,随机化和t检验,这与预期的错误率更一致。虽然统计检验在较小样本量下的功效存在差异,但在大样本量下的功效没有差异。我们建议不要使用符号和Wilcoxon检验来分析大规模评估结果。我们的结果表明,使用Top-N推荐和大搜索评估数据,大多数测试将有100%的机会找到统计上显着的结果。因此,应使用效应量来确定实际或科学意义。
A number of information retrieval studies have been done to assess which statistical techniques are appropriate for comparing systems. However, these studies are focused on TREC-style experiments, which typically have fewer than 100 topics. There is no similar line of work for large search and recommendation experiments; such studies typically have thousands of topics or users and much sparser relevance judgements, so it is not clear if recommendations for analyzing traditional TREC experiments apply to these settings. In this paper, we empirically study the behavior of significance tests with large search and recommendation evaluation data. Our results show that the Wilcoxon and Sign tests show significantly higher Type-1 error rates for large sample sizes than the bootstrap, randomization and t-tests, which were more consistent with the expected error rate. While the statistical tests displayed differences in their power for smaller sample sizes, they showed no difference in their power for large sample sizes. We recommend the sign and Wilcoxon tests should not be used to analyze large scale evaluation results. Our result demonstrate that with Top-N recommendation and large search evaluation data, most tests would have a 100% chance of finding statistically significant results. Therefore, the effect size should be used to determine practical or scientific significance.