TREC CAR Y3: Complex Answer Retrieval Overview

TREC CAR Y3: Complex Answer Retrieval Overview
复制标题

TREC CAR Y3:复杂答案检索概述

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Laura Dietz
Laura Dietz
中科院分区:
--
文献类型:
--
作者:
Laura Dietz

文献摘要

参考文献

被引文献

相似文献

TREC复杂答案检索的愿景是创建复杂的长格式答案,以响应各种各样的信息需求。一般来说,我们希望创造的答案让人想起维基百科文章或学校教科书(例如TQA)。然而,虽然维基百科的绝大多数文章都是关于人的,但在TREC CAR中,我们的目标是那些不太常见的信息需求,涵盖了流行科学、技术和疾病的主题。在TREC CAR的头两年,我们的目标是复制维基百科的文章。这提供了一个非常大规模的自动化基准,对神经排序研究[7,6]以及基于特征的排序模型[1,4]具有重要影响。缺点是我们必须禁止访问维基百科,可能包括维基百科的集合(例如ClueWeb),以及来自维基百科的知识库(我们提供了维基百科的一部分,可以从中构建一个排除基准主题的知识图,称为“allButBenchmark”)。从技术上讲,这甚至会影响到在维基百科上训练的资源,比如大多数词嵌入和BERT b[3]。为了避免这个困难,今年,我们的测试主题完全来自学校教科书章节的集合,这些章节与TQA数据集[5]一起提供。这些章节的长度与维基百科的文章相似,但是为年轻读者写的。我们从TQA章节中获得大纲,并要求参与者使用往年的段落语料库,用维基百科中的段落填充这些大纲。我们手动清理和重写了大纲,使它们适合被视为搜索查询。这个测试集合称为benchmarkY3test。这个决定的一个缺点是不能进行自动评估。我们建议在第一年(Y1)提供的“训练”集合上训练需要大量数据的方法。由于去年的测试数据(benchmarkY2test)同时包含了维基百科主题和TQA主题,我们重新发布了人工评估的TQA主题作为今年的训练数据,发布为benchmarkY3train。
The vision of TREC Complex Answer Retrieval is to create complex long-form answers in response to a wide-variety information needs. In general, we aspire to create answers that are reminiscent to Wikipedia articles or school text books (e.g. TQA). However, while the vast majority of Wikipedia articles are about people, in TREC CAR we aim at information needs that are off the beaten path, covering topics in popular science, technology, and illnesses. The first two years of TREC CAR, we aimed to reproduce Wikipedia articles. This provided a very large-scale automated benchmark, which had significant impact on neural ranking research [7, 6], as well as feature-based ranking models [1, 4]. The downside was that we had to prohibit access to Wikipedia, collections that could include Wikipedia (e.g. ClueWeb), and knowledge bases derived from Wikipedia (we provided the part of Wikipedia from which a knowledge graph can be built that excludes the benchmark topics, called “allButBenchmark”). Technically this would even affect resources that are trained onWikipedia, such as most word embeddings and BERT [3]. To avoid this difficulty, in this year, our test topics come exclusively from an collection of school text book chapters, which are provided along with the TQA dataset [5]. These chapters have a similar length as Wikipedia articles, but are written for a younger audience. We derive outlines from TQA chapters, and ask participants to populate these outlines with paragraphs from Wikipedia, using the paragraphCorpus from previous years. We manually cleaned and rewrote the outlines so that they are suitable to be treated like search queries. This test collection is called benchmarkY3test. A downside of this decision is that no automatic evaluation can be conducted. We recommend to train data-hungry methods on the “train” collection provided in the first year (Y1). Since previous year’s test data (benchmarkY2test ) contained both contained Wikipedia topics and TQA topics, we re-released manually assessed TQA topics as training data for this year, released as benchmarkY3train.
DOI: 10.1145/3331184.3331257
发表时间: 2019-07
期刊: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子: --
作者:
Laura Dietz
通讯作者: Laura Dietz