A publicly available benchmark for biomedical dataset retrieval: the reference standard for the 2016 bioCADDIE dataset retrieval challenge.

A publicly available benchmark for biomedical dataset retrieval: the reference standard for the 2016 bioCADDIE dataset retrieval challenge.
复制标题

DOI:
10.1093/database/bax061
复制
发表时间:
2017-01-01
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Xu H
Xu H
中科院分区:
其他
文献类型:
--
作者:
Cohen T;Roberts K;Gururaj AE;Chen X;Pournejati S;Alter G;Hersh WR;Demner-Fushman D;Ohno-Machado L;Xu H

文献摘要

参考文献

被引文献

相似文献

公开可用的生物医学数据集的快速增长提供了丰富的资源,这些资源可能具有复制先前实验以及生成和探索新假设的价值。然而,这些数据集分布在广泛的数据集储存库中,侧重于不同的数据类型,并使用不同的术语编制索引,因此在重新使用这些数据集方面存在一些障碍。需要新的方法,使生物医学研究人员能够在这个迅速扩大的信息生态系统中找到感兴趣的数据集,并需要新的资源来正式评估这些方法。在本文中,我们描述了生物医学数据集信息检索基准的设计和生成,该基准已被开发并用于2016年bioCADDIE数据集检索挑战赛。在传统的开创性的克兰菲尔德实验,并作为例证的文本检索会议(TREC),这个基准包括一个语料库(生物医学数据集),一组查询,和相关性判断这些查询的语料库的元素。本文介绍了这些元素的过程中,通过每个派生,重点是这些方面,区分这个基准从典型的信息检索参考集。具体而言,我们讨论了我们的查询的起源在一个更大的协作努力的背景下,生物医学和healthCAre数据发现指数生态系统(bioCADDIE)财团,以及生物医学数据集检索的显着特点作为一项任务。由此产生的基准集已公开提供,以推进生物医学数据集检索领域的研究。 数据库URL:https://biocaddie.org/benchmark-data
The rapid proliferation of publicly available biomedical datasets has provided abundant resources that are potentially of value as a means to reproduce prior experiments, and to generate and explore novel hypotheses. However, there are a number of barriers to the re-use of such datasets, which are distributed across a broad array of dataset repositories, focusing on different data types and indexed using different terminologies. New methods are needed to enable biomedical researchers to locate datasets of interest within this rapidly expanding information ecosystem, and new resources are needed for the formal evaluation of these methods as they emerge. In this paper, we describe the design and generation of a benchmark for information retrieval of biomedical datasets, which was developed and used for the 2016 bioCADDIE Dataset Retrieval Challenge. In the tradition of the seminal Cranfield experiments, and as exemplified by the Text Retrieval Conference (TREC), this benchmark includes a corpus (biomedical datasets), a set of queries, and relevance judgments relating these queries to elements of the corpus. This paper describes the process through which each of these elements was derived, with a focus on those aspects that distinguish this benchmark from typical information retrieval reference sets. Specifically, we discuss the origin of our queries in the context of a larger collaborative effort, the biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium, and the distinguishing features of biomedical dataset retrieval as a task. The resulting benchmark set has been made publicly available to advance research in the area of biomedical dataset retrieval. Database URL: https://biocaddie.org/benchmark-data
DOI: 10.1145/361219.361220
发表时间: 1975-01-01
影响因子: 22.7
作者:
SALTON, G;WONG, A;YANG, CS
通讯作者: YANG, CS
DOI: 10.1007/s10791-015-9259-x
发表时间: 2016-04-01
影响因子: 2.5
作者:
Roberts, Kirk;Simpson, Matthew;Hersh, William
通讯作者: Hersh, William
DOI: 10.1093/nar/gkp440
发表时间: 2009-07
影响因子: 14.9
作者:
Noy NF;Shah NH;Whetzel PL;Dai B;Dorf M;Griffith N;Jonquet C;Rubin DL;Storey MA;Chute CG;Musen MA
通讯作者: Musen MA
DOI: 10.1007/s10791-008-9072-x
发表时间: 2009-02-01
期刊: INFORMATION RETRIEVAL
影响因子: --
作者:
Roberts, Phoebe M.;Cohen, Aaron M.;Hersh, William R.
通讯作者: Hersh, William R.
DOI: 10.1038/505612a
发表时间: 2014-01-30
期刊: NATURE
影响因子: 64.8
作者:
Collins, Francis S.;Tabak, Lawrence A.
通讯作者: Tabak, Lawrence A.