Wikimarks: Harvesting Relevance Benchmarks from Wikipedia

Wikimarks: Harvesting Relevance Benchmarks from Wikipedia
复制标题

DOI:
10.1145/3477495.3531731
复制
发表时间:
2022-07
期刊:
Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
Laura Dietz;Shubham Chatterjee;Connor Lennox;Sumanta Kashyapi;P. Oza;Ben Gamari
Laura Dietz;Shubham Chatterjee;Connor Lennox;Sumanta Kashyapi;P. Oza;Ben Gamari
中科院分区:
其他
文献类型:
--
作者:
Laura Dietz;Shubham Chatterjee;Connor Lennox;Sumanta Kashyapi;P. Oza;Ben Gamari

文献摘要

相似文献

我们提供了从维基百科自动获取相关性基准的资源——我们将其称为“维基标记”,以将它们与手动创建的基准区分开来。与模拟基准不同,它们基于维基百科作者的手动注释。对 TREC 复杂答案检索轨道的研究表明,Wikimarks 下的排行榜和手动注释的基准非常相似。由于其可用性,维基标记可以满足信息检索研究的重要需求。我们提供元资源来收集维基标记,用于跨不同语言的多种信息检索任务:段落检索、实体排名、特定于查询的聚类、大纲预测和相关实体链接等等。此外,我们还提供源自 01/01/2022 维基百科转储的英语、简单英语和日语维基标记示例。可用资源:https://trema-unh.github.io/wikimarks/
We provide a resource for automatically harvesting relevance benchmarks from Wikipedia -- which we refer to as "Wikimarks" to differentiate them from manually created benchmarks. Unlike simulated benchmarks, they are based on manual annotations of Wikipedia authors. Studies on the TREC Complex Answer Retrieval track demonstrated that leaderboards under Wikimarks and manually annotated benchmarks are very similar. Because of their availability, Wikimarks can fill an important need for Information Retrieval research. We provide a meta-resource to harvest Wikimarks for several information retrieval tasks across different languages: paragraph retrieval, entity ranking, query-specific clustering, outline prediction, and relevant entity linking and many more. In addition, we provide example Wikimarks for English, Simple English, and Japanese derived from the 01/01/2022 Wikipedia dump. Resource available: https://trema-unh.github.io/wikimarks/