DataMed - an open source discovery index for finding biomedical datasets.

DataMed - an open source discovery index for finding biomedical datasets.
复制标题

DOI:
10.1093/jamia/ocx121
复制
发表时间:
2018-03-01
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Xu H
Xu H
中科院分区:
其他
文献类型:
--
作者:
Chen X;Gururaj AE;Ozyurt B;Liu R;Soysal E;Cohen T;Tiryaki F;Li Y;Zong N;Jiang M;Rogith D;Salimi M;Kim HE;Rocca-Serra P;Gonzalez-Beltran A;Farcas C;Johnson T;Margolis R;Alter G;Sansone SA;Fore IM;Ohno-Machado L;Grethe JS;Xu H

文献摘要

参考文献

被引文献

相似文献

寻找相关数据集对于促进生物医学领域的数据重用非常重要,但考虑到生物医学数据的数量和复杂性,这具有挑战性。在这里,我们描述了一个名为 DataMed 的开源生物医学数据发现系统的开发,其目标是促进生物医学领域附加数据索引的构建。 DataMed 是由美国国立卫生研究院资助的生物医学和 healthCAre 数据发现索引生态系统 (bioCADDIE) 联盟开发的,可以跨存储库有效地索引和搜索不同类型的生物医学数据集。它由 2 个主要组件组成:(1) 数据摄取管道,用于收集原始元数据信息并将其转换为统一的元数据模型,称为数据标签套件 (DATS);(2) 搜索引擎,用于根据用户输入的查询查找相关数据集。除了描述其架构和技术之外,我们还评估了 DataMed 中的各个组件,包括摄取管道的准确性、跨存储库的 DATS 模型的流行度以及数据集检索引擎的整体性能。我们的手动审查表明,摄取管道可以达到 90% 的准确性,并且 DATS 的核心元素在各个存储库中的频率各不相同。在手动策划的基准数据集上,DataMed 搜索引擎通过实施先进的自然语言处理和术语服务,推断平均精度为 0.2033,10 精度(P@10,前 10 个搜索结果中的相关结果数量)为 0.6022。目前,我们已将 DataMed 系统作为开源包向生物医学界公开。
Finding relevant datasets is important for promoting data reuse in the biomedical domain, but it is challenging given the volume and complexity of biomedical data. Here we describe the development of an open source biomedical data discovery system called DataMed, with the goal of promoting the building of additional data indexes in the biomedical domain. DataMed, which can efficiently index and search diverse types of biomedical datasets across repositories, is developed through the National Institutes of Health–funded biomedical and healthCAre Data Discovery Index Ecosystem (bioCADDIE) consortium. It consists of 2 main components: (1) a data ingestion pipeline that collects and transforms original metadata information to a unified metadata model, called DatA Tag Suite (DATS), and (2) a search engine that finds relevant datasets based on user-entered queries. In addition to describing its architecture and techniques, we evaluated individual components within DataMed, including the accuracy of the ingestion pipeline, the prevalence of the DATS model across repositories, and the overall performance of the dataset retrieval engine. Our manual review shows that the ingestion pipeline could achieve an accuracy of 90% and core elements of DATS had varied frequency across repositories. On a manually curated benchmark dataset, the DataMed search engine achieved an inferred average precision of 0.2033 and a precision at 10 (P@10, the number of relevant results in the top 10 search results) of 0.6022, by implementing advanced natural language processing and terminology services. Currently, we have made the DataMed system publically available as an open source package for the biomedical community.
DOI: 10.1016/j.jbi.2009.02.002
发表时间: 2009-04
影响因子: 4.5
作者:
Cohen T;Widdows D
通讯作者: Widdows D
DOI: 10.1186/1471-2105-11-255
发表时间: 2010-05-17
期刊: BMC bioinformatics
影响因子: 3
作者:
Chen B;Dong X;Jiao D;Wang H;Zhu Q;Ding Y;Wild DJ
通讯作者: Wild DJ
DOI: 10.1093/nar/gkr1178
发表时间: 2012-01
影响因子: 14.9
作者:
Federhen S
通讯作者: Federhen S
DOI: 10.1038/nbt1150
发表时间: 2006-01-01
影响因子: 46.9
作者:
Butte, AJ;Kohane, IS
通讯作者: Kohane, IS
DOI: 10.1093/nar/gkh036
发表时间: 2004-01-01
影响因子: 14.9
作者:
Harris, MA;Clark, J;White, R
通讯作者: White, R