The Metadata Powerwash - Integrated tools to make biomedical data FAIR
The Metadata Powerwash - Integrated tools to make biomedical data FAIR
批准号:
10093841
负责人:
Mark A Musen
金额:
$33.48万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-05-01 至 2025-01-31
关键词:
AgeBiological Specimen BanksCategoriesCollectionCommon Data ElementCommunitiesComputersDataData ScienceData SetDiseaseFAIR principlesFunding AgencyGoalsGoldInformation TechnologyKnowledgeLibrariesLinkManualsMetadataMethodsNamesNatural Language ProcessingNumerical valueOntologyPharmaceutical PreparationsProblem SolvingProcessRecordsReportingReproducibilityResearchResearch PersonnelResourcesRetrievalSamplingScienceScientistSpecific qualifier valueSpeedStandardizationStructureTechnologyTestingTimeVariantWorkbasebiomedical scientistdata archivedata repositorydata reuseexperimental studyimprovedindexinginformation organizationinteroperabilitymetadata standardspublic repositoryrepositorysample collectionsearch enginesecondary analysistool
中文摘要
项目摘要
描述科学数据的元数据是实现(1)
数据的发现和再利用以及(2)产生数据的实验的可重复性
数据放在首位。元数据对于科学家理解相关数据至关重要
并重复使用它们,以及信息技术对数据进行索引,使数据
可用,并为科学家搜索相应的数据集提供过滤器。
目前,公共存储库中托管的科学元数据存在多个质量问题
这限制了科学家发现和重复使用他们所指的实验数据集的能力。它可以
花费科学家数周的时间来确定满足特定要求的数据集的集合
当数据描述得如此之差时的标准-并且过程的大部分必然
手册。
我们建议开发一个端到端的解决方案来标准化生物医学元数据
本体的帮助-定义应用程序域中的术语的数据结构
他们之间的关系。有数百种本体论为以下内容提供标准术语
在生物医学中使用,它们是制作生物医学元数据的必要资源
可互操作和可重复使用。我们的方法还将建立在
扩展数据注释和检索中心(Cedar),该中心提供
用于定义基于计算机的元数据模板的块和公共数据元素
社区标准。
我们的计划涉及三个具体目标。首先,我们将开发一种方法和工具来标准化
可能出现在元数据中以表示相同的多个即席元数据字段名称
通过将这些字段名替换为标准中使用的字段名来确定信息类型
元数据模板,或者,如果没有合适的模板匹配,则使用相关
本体论。其次,我们将开发方法和工具来标准化不同类型的元数据
字段值,例如,药品或疾病等类型值和数值
例如年龄或样本采集日期。第三,我们将评估速度、精确度和召回率
我们的元数据转换管道-由标准化字段的方法和工具构建而成
名称和值-基于我们将手动管理的大型元数据语料库
现有公共元数据。我们还将进行实验,以测试
当生物医学科学家在其上下文中执行数据集搜索时标准化的元数据
工作。
英文摘要
Project Summary
The metadata that describe scientific data are fundamental resources to enable (1) the
discovery and reuse of the data and (2) the reproducibility of the experiments that generated the
data in the first place. Metadata are essential for scientists to understand the associated data
and to reuse them, as well as for information technology to index the data, to make the data
available, and to provide filters for scientists to search for the corresponding datasets.
Currently, the scientific metadata hosted in public repositories suffer from multiple quality issues
that limit scientists’ ability to find and reuse the experimental datasets to which they refer. It can
take many weeks of a scientist’s time to identify a collection of datasets that fulfill specific
criteria when the data are so poorly described—and the majority of the process is necessarily
manual.
We propose to develop an end-to-end solution to standardize biomedical metadata with the
help of ontologies—data structures that define the terms in an application domain and the
relationships among them. There are hundreds of ontologies that provide standard terms for
use in biomedicine, and they are essential resources to make biomedical metadata
interoperable and reusable. Our approach also will build on the technology created by the
Center for Expanded Data Annotation and Retrieval (CEDAR), which offers a library of building
blocks and common data elements for defining computer-based metadata templates based on
community standards.
Our plan involves three specific aims. First, we will develop a method and tool to standardize
the multiple, ad hoc metadata field names that may appear in metadata to represent the same
type of information by replacing those field names with the field names used in standard
metadata templates or, if no appropriate template match is available, with terms from a relevant
ontology. Second, we will develop methods and tools to standardize different types of metadata
field values, for example, categorical values such as drugs or diseases, and numerical values
such as age, or sample collection date. Third, we will evaluate the speed, precision, and recall
of our metadata transformation pipeline—built out of the methods and tools to standardize field
names and values—on a large corpus of metadata that we will manually curate based on
existing public metadata. We will also carry out experiments to test the effect of the
standardized metadata when biomedical scientists perform dataset search in the context of their
work.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enhanced ontology engineering through a Web-based, Cloud-based software architecture
-
批准号:10405968
-
项目类别:
-
资助金额:$23.61万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
The Metadata Powerwash - Integrated tools to make biomedical data FAIR
-
批准号:10397981
-
项目类别:
-
资助金额:$33.45万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Enhancing the RADx Data Hub for Data FAIRness
-
批准号:10433797
-
项目类别:
-
资助金额:$300.0万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Enhancing the RADx Data Hub for Data FAIRness
-
批准号:10794704
-
项目类别:
-
资助金额:$1010.0万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Improved metadata authoring to enhance AI/ML readiness of associated datasets
-
批准号:10592638
-
项目类别:
-
资助金额:$27.45万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
The Metadata Powerwash - Integrated tools to make biomedical data FAIR
-
批准号:10551273
-
项目类别:
-
资助金额:$33.45万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
BioPortal: An Expansive Knowledgebase of Biomedical Entities and Relations
-
批准号:10494104
-
项目类别:
-
资助金额:$107.39万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
BioPortal: An Expansive Knowledgebase of Biomedical Entities and Relations
-
批准号:10271048
-
项目类别:
-
资助金额:$108.3万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Enhancing the RADx Data Hub for Data FAIRness
-
批准号:10699372
-
项目类别:
-
资助金额:$23.16万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Enhancing the RADx Data Hub for Data FAIRness
-
批准号:10850055
-
项目类别:
-
资助金额:$3100.0万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
A Knowledge Provider for Scruffy Sources of Metadata in Translational Medicine
-
批准号:10057243
-
项目类别:
-
资助金额:$5.6万
-
财政年份:2020
-
负责人:Mark A Musen
-
依托单位:
Center for Expanded Data Annotation and Retrieval (CEDAR) - Overall
-
批准号:8774419
-
项目类别:
-
资助金额:$185.47万
-
财政年份:2014
-
负责人:Mark A Musen
-
依托单位:
Center for Expanded Data Annotation and Retrieval (CEDAR) - Overall
-
批准号:9052149
-
项目类别:
-
资助金额:$359.88万
-
财政年份:2014
-
负责人:Mark A Musen
-
依托单位:
Center for Expanded Data Annotation and Retrieval (CEDAR) - Overall
-
批准号:8935754
-
项目类别:
-
资助金额:$329.99万
-
财政年份:2014
-
负责人:Mark A Musen
-
依托单位:
Protege: An Ontology-Development Platform for Biomedical Scientists
-
批准号:8438361
-
项目类别:
-
资助金额:$53.36万
-
财政年份:2013
-
负责人:Mark A Musen
-
依托单位:
Protege: An Ontology-Development Platform for Biomedical Scientists
-
批准号:8987580
-
项目类别:
-
资助金额:$52.65万
-
财政年份:2013
-
负责人:Mark A Musen
-
依托单位:
Protege: An Ontology-Development Platform for Biomedical Scientists
-
批准号:8597446
-
项目类别:
-
资助金额:$53.36万
-
财政年份:2013
-
负责人:Mark A Musen
-
依托单位:
Protege: An Ontology-Development Platform for Biomedical Scientists
-
批准号:8788417
-
项目类别:
-
资助金额:$52.65万
-
财政年份:2013
-
负责人:Mark A Musen
-
依托单位:
Core 4
-
批准号:8045892
-
项目类别:
-
资助金额:$44.0万
-
财政年份:2010
-
负责人:Mark A Musen
-
依托单位:
Core 2
-
批准号:8045887
-
项目类别:
-
资助金额:$27.99万
-
财政年份:2010
-
负责人:Mark A Musen
-
依托单位:
海外基金