Statistical Significance Testing in Theory and in Practice

Statistical Significance Testing in Theory and in Practice
复制标题

理论和实践中的统计显着性检验

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on the Theory of Information Retrieval
影响因子:
--
通讯作者:
Ben Carterette
Ben Carterette
中科院分区:
--
文献类型:
--
作者:
Ben Carterette

文献摘要

被引文献

相似文献

过去25年来,信息访问问题实验的严格性有了很大提高。这主要是由于三个因素:高质量,公共的,便携式的测试集合,如TREC(文本检索会议~citetrecbook)制作的测试集合,大量用户群体在线A/B测试的增加,以及统计假设检验的增加,以确定观察到的改善是否可以归因于随机机会之外的其他因素。这些共同为评审员、项目委员会和期刊编辑创造了一个非常有用的标准;关于信息访问(IA)问题的工作,如搜索和推荐,越来越不能发表,除非它已经使用一个构造良好的测试集合进行离线评估,或者在一个大的用户基础上进行在线评估,并显示在良好的基线上产生统计学上的显着改善。但是,正如谚语所说,任何锋利到足以有用的工具也锋利到足以危险。显著性的统计检验被广泛误解。大多数研究人员和开发人员将其视为一个“黑盒子”:评估结果进去,p值出来。但是,由于重要性是决定探索哪些方向以及发布或部署哪些内容的重要因素,因此使用未经思考获得的p值可能会对IA中的每个人产生影响。约安纳蒂认为,生物医学科学的主要后果是,大多数已发表的研究结果是错误的;这是否也适用于IA?
The past 25 years have seen a great improvement in the rigor of experimentation on information access problems. This is due primarily to three factors: high-quality, public, portable test collections such as those produced by TREC (the Text REtreval Conference~citetrecbook ), the increased ease of online A/B testing on large user populations, and the increased practice of statistical hypothesis testing to determine whether observed improvements can be ascribed to something other than random chance. Together these create a very useful standard for reviewers, program committees, and journal editors; work on information access (IA) problems such as search and recommendation increasingly cannot be published unless it has been evaluated offline using a well-constructed test collection or online on a large user base and shown to produce a statistically significant improvement over a good baseline. But, as the saying goes, any tool sharp enough to be useful is also sharp enough to be dangerous. Statistical tests of significance are widely misunderstood. Most researchers and developers treat them as a "black box': evaluation results go in and a p-value comes out. But because significance is such an important factor in determining what directions to explore and what is published or deployed, using p-values obtained without thought can have consequences for everyone working in IA. Ioannidis has argued that the main consequence in the biomedical sciences is that most published research findings are false; could that be the case for IA as well?