TripClick: The Log Files of a Large Health Web Search Engine

TripClick: The Log Files of a Large Health Web Search Engine
复制标题

DOI:
10.1145/3404835.3463242
复制
发表时间:
2021-03
期刊:
Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Navid Rekabsaz;Oleg Lesota;M. Schedl;J. Brassey;Carsten Eickhoff
Navid Rekabsaz;Oleg Lesota;M. Schedl;J. Brassey;Carsten Eickhoff
中科院分区:
其他
文献类型:
--
作者:
Navid Rekabsaz;Oleg Lesota;M. Schedl;J. Brassey;Carsten Eickhoff

文献摘要

被引文献

相似文献

点击日志是各种信息检索 (IR) 任务的宝贵资源。这包括查询理解/分析,以及学习有效的 IR 模型,特别是当模型需要大量训练数据时。我们发布了一个大规模的特定领域点击日志数据集,这些数据集是从旅行数据库健康网络搜索引擎的用户交互中获得的。我们的点击日志数据集包含 2013 年至 2020 年间收集的约 520 万个用户交互。我们使用该数据集创建标准 IR 评估基准 - TripClick - 包含约 700,000 个独特的自由文本查询和 130 万对查询文档相关性信号,其相关性由两个点击模型估计。因此,该集合是为数不多的提供必要数据丰富性和规模的数据集之一,用于训练具有大量参数的神经 IR 模型,尤其是健康领域的第一个数据集。使用 TripClick,我们进行实验来评估各种 IR 模型,展示利用这些数据来训练神经架构的好处。特别是,评估结果表明,相对于经典 IR 模型,性能最佳的神经 IR 模型显着提高了性能,尤其是对于更频繁的查询。
Click logs are valuable resources for a variety of information retrieval (IR) tasks. This includes query understanding/analysis, as well as learning effective IR models particularly when the models require large amounts of training data. We release a large-scale domain-specific dataset of click logs, obtained from user interactions of the Trip Database health web search engine. Our click log dataset comprises approximately 5.2 million user interactions collected between 2013 and 2020. We use this dataset to create a standard IR evaluation benchmark - TripClick - with around 700,000 unique free-text queries and 1.3 million pairs of query-document relevance signals, whose relevance is estimated by two click-through models. As such, the collection is one of the few datasets offering the necessary data richness and scale to train neural IR models with a large amount of parameters, and notably the first in the health domain. Using TripClick, we conduct experiments to evaluate a variety of IR models, showing the benefits of exploiting this data to train neural architectures. In particular, the evaluation results show that the best performing neural IR model significantly improves the performance by a large margin relative to classical IR models, especially for more frequent queries.