Challenges of Multileaved Comparison in Practice: Lessons from NTCIR-13 OpenLiveQ Task.
Challenges of Multileaved Comparison in Practice: Lessons from NTCIR-13 OpenLiveQ Task.
复制标题
实践中多叶比较的挑战:NTCIR-13 OpenLiveQ 任务的经验教训。
DOI:
10.1145/3269206.3269318
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
Makoto P. Kato,Tomohiro Manabe,Sumio Fujita,Akiomi Nishida,Takehiro Yamamoto
中科院分区:
文献类型:
--
作者:
Suppanut Pothirattanachaikul;Takehiro Yamamoto;Sumio Fujita;Akira Tajima;Katsumi Tanaka;Masatoshi Yoshikawa;小國宏樹,児玉哲司,川﨑忠寛,生田孝;Makoto P. Kato,Tomohiro Manabe,Sumio Fujita,Akiomi Nishida,Takehiro Yamamoto
This paper discusses challenges of an online evaluation technique, multileaved comparison, based on the analysis of evaluation results in a community question-answering (cQA) search service. NTCIR-13 OpenLiveQ task offered a shared task in which participants addressed an ad-hoc retrieval task in a cQA service, and evaluated their rankers by multileaved comparison, which combines multiple rankings to generate a single search result page, and simultaneously evaluates the different rankings based on users' clicks on the search result page. Since the number of search result impressions during the evaluation period might not suffice to evaluate a hundred of rankers, we conducted the online evaluation only for rankers that achieved high performance in offline evaluation. The analysis of evaluation results showed that offline and online evaluation results did not fully agree, and a large number of users' clicks were necessary to find a statistically significant difference for every ranker pair. To cope with these problems in large-scale multileaved comparison, we propose a new experimental design that evaluates all the rankers online but intensively tests only the top-k rankers. Simulation-based experiments demonstrated that Copeland counting algorithm could achieve high top-k recall in the top-k identification problem for multileaved comparison.