A Knowledge Provider for Scruffy Sources of Metadata in Translational Medicine
A Knowledge Provider for Scruffy Sources of Metadata in Translational Medicine
批准号:
10057243
负责人:
Mark A Musen
金额:
$5.6万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-01-23 至 2020-04-07
关键词:
AddressAreaBig Data to KnowledgeBiomedical ComputingClinical TrialsCommunitiesComputer softwareComputersDataData CollectionData SetData SourcesEvaluationGoalsKnowledgeLaboratoriesLiteratureManualsMetadataMethodologyMethodsOntologyPatientsPeer ReviewPerformanceProcessProviderPublicationsPublishingRecordsResearchResourcesScientistSemanticsServicesSourceStandardizationSystemTechniquesTechnologyTestingTextUnited States National Institutes of HealthWorkbiomedical ontologydata dictionarydata standardsexperimental studyinterestinteroperabilityknowledge graphknowledge integrationonline repositorypatient populationprogramsprospectiverepositoryresponsesecondary analysisspellingstatisticstranslational medicine
中文摘要
生物医学数据翻译器的一项基本任务是识别
已经执行或正在进行的,并能够整合关于
实验方法、结果,以及--如果有的话--与其他知识的结论
消息来源。这样的功能将支持以下查询:(1)是否有人执行过
使用这样的方法进行实验?(2)有人进行过这样的研究吗?
支持某一特定结论?(3)是否有针对某一特定疾病的临床试验
病人群适合我现在需要治疗的病人吗?(4)什么是最好的?
针对某一特定情况的当前临床试验的结果是否建议进行实践?
有时,这些疑问可以通过对科学文献的分析来解决。更多
然而,出版的文献往往没有提供所需的方法论细节
解决这样的问题--即使NLP技术足够好,可以找到答案。
出版物也只提供了实验结果的汇总统计数据。要解决这个问题
翻译人员最感兴趣的查询类型,有必要访问实际的
在线实验数据,从旨在提供以下描述的元数据开始
数据集和最初导致数据收集的实验的数据。
翻译器项目的问题是,描述大多数在线数据的元数据
计算机很难找到和处理实验数据源。我们实验室的
例如,对NCBI生物样本元数据存储库的分析表明,科学家在很大程度上
避免完全使用标准数据字典,而且--部分结果是--它们非常
当他们提供元数据值时草率[3]。(一个恰当的例子:大约76%的元数据
BiosSample中应为布尔值的值既不为真也不为假。)尽管有这么多
在过去几年中关于使在线数据集可查找、可访问、
可互操作、可重复使用的(公平)[14],大多数在线数据集并不接近公平。
我们的实验室正在开发可以纠正在线元数据错误的技术。就像一个
元数据的拼写检查器,我们的方法将尝试识别元数据的意图
作者,以纠正打字错误,并尽可能将自由文本字符串转换为本体术语
[6]。我们的目标是提供一种服务,将网上弥漫的杂乱无章的元数据
将生物医学实验的描述转化为允许自动发现的形式,
以一种根本不可能的方式整合和二次分析研究结果
现在时。我们预计翻译者将调用我们的服务来查找实验数据集
及其随附的元数据,以执行此类数据集的标准分析,并
将对实验的描述整合到不断演变的知识图谱中。
我们将通过研究知识提供商对查询的响应来评估其性能
来自翻译社区,并通过同行审查基础的子集,清理
它从实际的在线存储库处理的元数据记录,如BiosSample和
临床试验.gov。我们的评估必然会受到选择一个
元数据测试集的可管理性和人工同行评审的固有缺陷。
我们的实验室有合作开发国家重大资源的持久传统
将语义技术引入生物医学。我们的BioPortal本体库[5]是
由国家生物医学本体中心(NCBO)开发,NIH国家生物医学本体论中心之一
生物医学计算中心。面向未来创作的Cedar工作台
标准化元数据[11,12]是在NIH大数据到知识(BD2K)下开发的
程序。我们用于构建和维护生物医学本体的Protégé系统是
世界上广泛使用的创建语义技术的软件[15]。我们的小组正在进行
与Pinterest、BASF和Elsevier等公司建立关系,以帮助他们
努力开发企业范围的知识图谱。因此,我们完全有能力发展我们的
知识提供者,并在语义技术领域广泛协助该财团。
英文摘要
An essential task for the Biomedical Data Translator is to identify scientific experiments that
have been performed or that are ongoing, and to enable integration of knowledge of the
experimental methods, the results, and—when available—the conclusions with other knowledge
sources. Such capabilities will enable queries such as: (1) Has anyone ever performed an
experiment using methods like these? (2) Has anyone performed a study where the data may
support a particular conclusion? (3) Are there any clinical trials for a particular condition whose
patient population is a good match for a patient whom I now need to treat? (4) What best
practices are suggested by the results of current clinical trials for a particular condition?
Sometimes such queries can be addressed through an analysis of the scientific literature. More
often, however, the published literature does not provide the methodological details needed to
address such questions—even if NLP techniques were good enough to find the answers.
Publications also provide only summary statistics of the experimental results. To address the
kinds of queries that are of most interest to the Translator, it is necessary to access the actual
experimental data online, starting with the metadata that are intended to provide descriptions of
the datasets and of the experiments that led to the collection of the data in the first place.
The problem for the Translator project is that the metadata that describe most online
experimental data sources are difficult for computers to find and to process. Our laboratory’s
analysis of the NCBI BioSample metadata repository, for example, shows that scientists largely
avoid using standard data dictionaries entirely, and—partly as a result—they are extremely
sloppy when they provide metadata values [3]. (A case in point: Some 76% of the metadata
values in BioSample that are intended to be Boolean are neither true nor false.) Despite all the
discussion in the past few years about making online datasets Findable, Accessible,
Interoperable, and Re-usable (FAIR) [14], most online datasets are not close to FAIR.
Our laboratory is developing technology that can rectify errors in online metadata. Like a
spell-checker for metadata, our approach will attempt to identify the intentions of metadata
authors, to correct typos, and to convert free-text strings to ontology terms whenever possible
[6]. Our goal is to provide a service that will transform the scruffy metadata that pervade online
descriptions of biomedical experiments into a form that will allow automated discovery,
integration, and secondary analysis of research results in ways that are simply not possible at
present. We anticipate that the Translator will call on our service to find experimental datasets
and their accompanying metadata, to perform standard analyses of such datasets, and to
integrate descriptions of experiments into the evolving knowledge graph.
We will evaluate the performance of our Knowledge Provider by studying its response to queries
from the Translator community and by peer review of a subset of the underlying, cleaned up
metadata records that it processes from actual online repositories, such as BioSample and
ClinicalTrials.gov. Our evaluation necessarily will be limited by the pragmatics of selecting a
manageable test set of metadata and by the inherent shortcomings of manual peer review.
Our laboratory has a sustained tradition of collaborating to develop major national resources
that bring semantic technology to biomedicine. Our BioPortal ontology repository [5] was
developed by the National Center for Biomedical Ontology (NCBO), one of the NIH National
Centers for Biomedical Computing. The CEDAR Workbench for the prospective authoring of
standardized metadata [11,12] was developed under the NIH Big Data to Knowledge (BD2K)
program. Our Protégé system for building and maintaining biomedical ontologies is the most
widely used software for creating semantic technology in the world [15]. Our group has ongoing
relationships with corporations such as Pinterest, BASF, and Elsevier to assist them in their
work to develop enterprise-wide knowledge graphs. We are thus well equipped to develop our
Knowledge Provider and to assist the consortium broadly in the area of semantic technology.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enhanced ontology engineering through a Web-based, Cloud-based software architecture
-
批准号:10405968
-
项目类别:
-
资助金额:$23.61万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
The Metadata Powerwash - Integrated tools to make biomedical data FAIR
-
批准号:10397981
-
项目类别:
-
资助金额:$33.45万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Enhancing the RADx Data Hub for Data FAIRness
-
批准号:10433797
-
项目类别:
-
资助金额:$300.0万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Enhancing the RADx Data Hub for Data FAIRness
-
批准号:10794704
-
项目类别:
-
资助金额:$1010.0万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Improved metadata authoring to enhance AI/ML readiness of associated datasets
-
批准号:10592638
-
项目类别:
-
资助金额:$27.45万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
The Metadata Powerwash - Integrated tools to make biomedical data FAIR
-
批准号:10551273
-
项目类别:
-
资助金额:$33.45万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
BioPortal: An Expansive Knowledgebase of Biomedical Entities and Relations
-
批准号:10494104
-
项目类别:
-
资助金额:$107.39万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
BioPortal: An Expansive Knowledgebase of Biomedical Entities and Relations
-
批准号:10271048
-
项目类别:
-
资助金额:$108.3万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Enhancing the RADx Data Hub for Data FAIRness
-
批准号:10699372
-
项目类别:
-
资助金额:$23.16万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
The Metadata Powerwash - Integrated tools to make biomedical data FAIR
-
批准号:10093841
-
项目类别:
-
资助金额:$33.48万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Enhancing the RADx Data Hub for Data FAIRness
-
批准号:10850055
-
项目类别:
-
资助金额:$3100.0万
-
财政年份:2021
-
负责人:Mark A Musen
-
依托单位:
Center for Expanded Data Annotation and Retrieval (CEDAR) - Overall
-
批准号:8774419
-
项目类别:
-
资助金额:$185.47万
-
财政年份:2014
-
负责人:Mark A Musen
-
依托单位:
Center for Expanded Data Annotation and Retrieval (CEDAR) - Overall
-
批准号:9052149
-
项目类别:
-
资助金额:$359.88万
-
财政年份:2014
-
负责人:Mark A Musen
-
依托单位:
Center for Expanded Data Annotation and Retrieval (CEDAR) - Overall
-
批准号:8935754
-
项目类别:
-
资助金额:$329.99万
-
财政年份:2014
-
负责人:Mark A Musen
-
依托单位:
Protege: An Ontology-Development Platform for Biomedical Scientists
-
批准号:8438361
-
项目类别:
-
资助金额:$53.36万
-
财政年份:2013
-
负责人:Mark A Musen
-
依托单位:
Protege: An Ontology-Development Platform for Biomedical Scientists
-
批准号:8987580
-
项目类别:
-
资助金额:$52.65万
-
财政年份:2013
-
负责人:Mark A Musen
-
依托单位:
Protege: An Ontology-Development Platform for Biomedical Scientists
-
批准号:8597446
-
项目类别:
-
资助金额:$53.36万
-
财政年份:2013
-
负责人:Mark A Musen
-
依托单位:
Protege: An Ontology-Development Platform for Biomedical Scientists
-
批准号:8788417
-
项目类别:
-
资助金额:$52.65万
-
财政年份:2013
-
负责人:Mark A Musen
-
依托单位:
Core 4
-
批准号:8045892
-
项目类别:
-
资助金额:$44.0万
-
财政年份:2010
-
负责人:Mark A Musen
-
依托单位:
Core 2
-
批准号:8045887
-
项目类别:
-
资助金额:$27.99万
-
财政年份:2010
-
负责人:Mark A Musen
-
依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
-
批准号:2021JJ40433
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2021
-
负责人:孙磊
-
依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
-
批准号:32001603
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:段真珍
-
依托单位:
AREA国际经济模型的移植.改进和应用
-
批准号:18870435
-
项目类别:面上项目
-
资助金额:2.0万元
-
批准年份:1988
-
负责人:史树中
-
依托单位: