Designing Test Collections for Comparing Many Systems

Designing Test Collections for Comparing Many Systems
复制标题

设计用于比较许多系统的测试集

DOI:
10.1145/2661829.2661893
复制
发表时间:
2014
期刊:
Proceedings of the 23rd ACM International Conference on Conference on Information and Knowledge Management
影响因子:
--
通讯作者:
T. Sakai
T. Sakai
中科院分区:
--
文献类型:
--
作者:
T. Sakai

文献摘要

被引文献

相似文献

一位研究人员决定建立一个测试集,用于将她的新信息检索(IR)系统与几个最先进的基线进行比较。她想知道她需要提前创建的主题的数量(n),这样她就可以开始寻找(比如说)足够大的查询日志来采样n个好的主题,并估计相关性评估成本。我们提供了实用的解决方案,像她这样的研究人员使用功率分析和样本量设计技术,并证明其实用性的几个IR任务和评估措施。我们不仅考虑了配对t检验,还考虑了单因素方差分析(ANOVA)进行显著性检验,以适应在给定的一组统计要求(α:I型错误率,α:II型错误率,minD:最佳和最差系统之间的最小可检测差异)下对m(≥ 2)个系统进行比较。使用我们简单的Excel工具和一些来自过去数据的合并方差估计,研究人员可以设计出统计上设计良好的测试集。我们证明,由于不同的评估措施有不同的主题之间的差异,他们不可避免地需要不同的主题集的大小。这表明,应在测试集设计阶段选择评估措施。此外,通过池深度减少实验与过去的数据,我们展示了如何相关性评估成本可以显着降低,同时冻结的统计要求。基于成本分析和可用预算,研究人员可以确定n和池深度pd之间的正确平衡。我们的技术和工具也适用于非IR任务的测试集合。
A researcher decides to build a test collection for comparing her new information retrieval (IR) systems with several state-of-the-art baselines. She wants to know the number of topics (n) she needs to create in advance, so that she can start looking for (say) a query log large enough for sampling n good topics, and estimating the relevance assessment cost. We provide practical solutions to researchers like her using power analysis and sample size design techniques, and demonstrate its usefulness for several IR tasks and evaluation measures. We consider not only the paired t-test but also one-way analysis of variance (ANOVA) for significance testing to accommodate comparison of m(≥ 2) systems under a given set of statistical requirements (α: the Type I error rate, ß: the Type II error rate, and minD: the minimum detectable difference between the best and the worst systems). Using our simple Excel tools and some pooled variance estimates from past data, researchers can design statistically well-designed test collections. We demonstrate that, as different evaluation measures have different variances across topics, they inevitably require different topic set sizes. This suggests that the evaluation measures should be chosen at the test collection design phase. Moreover, through a pool depth reduction experiment with past data, we show how the relevance assessment cost can be reduced dramatically while freezing the set of statistical requirements. Based on the cost analysis and the available budget, researchers can determine the right balance between n and the pool depth pd. Our techniques and tools are applicable to test collections for non-IR tasks as well.