Wikimarks: Harvesting Relevance Benchmarks from Wikipedia
Wikimarks: Harvesting Relevance Benchmarks from Wikipedia
复制标题
DOI:
10.1145/3477495.3531731
复制
发表时间:
2022-07
期刊:
影响因子:
--
通讯作者:
Laura Dietz;Shubham Chatterjee;Connor Lennox;Sumanta Kashyapi;P. Oza;Ben Gamari
中科院分区:
文献类型:
--
作者:
Laura Dietz;Shubham Chatterjee;Connor Lennox;Sumanta Kashyapi;P. Oza;Ben Gamari
We provide a resource for automatically harvesting relevance benchmarks from Wikipedia -- which we refer to as "Wikimarks" to differentiate them from manually created benchmarks. Unlike simulated benchmarks, they are based on manual annotations of Wikipedia authors. Studies on the TREC Complex Answer Retrieval track demonstrated that leaderboards under Wikimarks and manually annotated benchmarks are very similar. Because of their availability, Wikimarks can fill an important need for Information Retrieval research. We provide a meta-resource to harvest Wikimarks for several information retrieval tasks across different languages: paragraph retrieval, entity ranking, query-specific clustering, outline prediction, and relevant entity linking and many more. In addition, we provide example Wikimarks for English, Simple English, and Japanese derived from the 01/01/2022 Wikipedia dump. Resource available: https://trema-unh.github.io/wikimarks/