Challenges of Multileaved Comparison in Practice: Lessons from NTCIR-13 OpenLiveQ Task.

Challenges of Multileaved Comparison in Practice: Lessons from NTCIR-13 OpenLiveQ Task.
复制标题

实践中多叶比较的挑战:NTCIR-13 OpenLiveQ 任务的经验教训。

DOI:
10.1145/3269206.3269318
复制
发表时间:
2018
期刊:
Proceedings of the 27th ACM International Conference on Information and Knowledge Management, CIKM 2018
影响因子:
--
通讯作者:
Makoto P. Kato,Tomohiro Manabe,Sumio Fujita,Akiomi Nishida,Takehiro Yamamoto
Makoto P. Kato,Tomohiro Manabe,Sumio Fujita,Akiomi Nishida,Takehiro Yamamoto
中科院分区:
--
文献类型:
--
作者:
Suppanut Pothirattanachaikul;Takehiro Yamamoto;Sumio Fujita;Akira Tajima;Katsumi Tanaka;Masatoshi Yoshikawa;小國宏樹,児玉哲司,川﨑忠寛,生田孝;Makoto P. Kato,Tomohiro Manabe,Sumio Fujita,Akiomi Nishida,Takehiro Yamamoto

文献摘要

相似文献

本文在对社区问答(cQA)搜索服务评价结果分析的基础上,讨论了在线评价技术——多叶比较所面临的挑战。ntcirr -13 OpenLiveQ任务提供了一个共享任务,参与者在cQA服务中处理一个特别的检索任务,并通过多叶比较来评估他们的排名,多叶比较将多个排名组合成一个单一的搜索结果页面,同时根据用户在搜索结果页面上的点击来评估不同的排名。由于评估期间的搜索结果展示次数可能不足以评估一百个排名者,因此我们只对在线下评估中取得较高表现的排名者进行在线评估。对评价结果的分析表明,线下和线上的评价结果并不完全一致,需要大量的用户点击才能发现每对排名对有统计学意义的差异。为了解决大规模多叶比较中的这些问题,我们提出了一种新的实验设计,在线评估所有排名者,但只集中测试前k名的排名者。基于仿真的实验表明,Copeland计数算法能够在多叶比较的top-k识别问题中获得较高的top-k召回率。
This paper discusses challenges of an online evaluation technique, multileaved comparison, based on the analysis of evaluation results in a community question-answering (cQA) search service. NTCIR-13 OpenLiveQ task offered a shared task in which participants addressed an ad-hoc retrieval task in a cQA service, and evaluated their rankers by multileaved comparison, which combines multiple rankings to generate a single search result page, and simultaneously evaluates the different rankings based on users' clicks on the search result page. Since the number of search result impressions during the evaluation period might not suffice to evaluate a hundred of rankers, we conducted the online evaluation only for rankers that achieved high performance in offline evaluation. The analysis of evaluation results showed that offline and online evaluation results did not fully agree, and a large number of users' clicks were necessary to find a statistically significant difference for every ranker pair. To cope with these problems in large-scale multileaved comparison, we propose a new experimental design that evaluates all the rankers online but intensively tests only the top-k rankers. Simulation-based experiments demonstrated that Copeland counting algorithm could achieve high top-k recall in the top-k identification problem for multileaved comparison.