CAREER: Advancing the Role of Ontologies for Data Science in Biomedicine
CAREER: Advancing the Role of Ontologies for Data Science in Biomedicine
批准号:
2047001
负责人:
Licong Cui
金额:
$53.35万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-01 至 2026-08-31
中文摘要
本体是知识领域中概念(或类)、属性和概念之间关系的形式化表示。本体和术语在生物医学研究中发挥了至关重要的作用,用于编码,管理,共享和交换大量不断生成的异构生物医学数据,例如电子健康记录(EHR)。EHR已广泛用于转化研究,以学习预测模型,用于不同患者队列的发现和疾病管理。这种基于EHR的应用的第一步通常涉及患者队列识别。队列识别涉及资格标准的集合的规范,其需要使用EHR的语义主干(即,编码系统或本体),然后才能对EHR数据库运行查询。然而,有两个关键的障碍,在执行有效的队列识别从大规模的EHR。第一个是数据(或语义)异构性,由编码系统的混合使用引起。第二个是质量的语义骨干或本体层次结构,这是必不可少的翻译患者的资格标准,可执行的数据库查询。为了应对这些挑战,该项目将开发本体匹配和本体质量增强的新方法,这些方法直接影响生物医学中的数据科学实践,例如患者队列识别。此外,本项目将把建议的计算方面纳入以数据科学为基础的课程,以培养下一代数据科学家。本项目包括三个研究目标。在目标1中,PI将开发新的基于图神经网络(GNN)的学习方法,通过利用统一医学语言系统等来源中嵌入的知识来匹配生物医学本体。这将解决异构性问题并实现语义互操作性。在目标2中,PI将开发基于学习的方法来检测子类关系中的质量缺陷。这将解决质量问题,并实现本体层次结构的持续增强。在目标3中,PI将开发一个基于本体的COVID-19查询引擎,用于患者队列识别,这是一个增强语义互操作性的实际应用,用于支持数据驱动的COVID-19研究。为了评估所提出的方法,领域专家将参与验证所产生的匹配概念和检测到的质量问题。PI将向各自的本体所有者传达经过验证的质量问题,以便在后续的本体版本中进行纠正。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
An ontology is a formal representation of concepts (or classes), properties, and relationships between concepts within a knowledge domain. Ontologies and terminologies have played a vital role in biomedical research for coding, managing, sharing, and exchange of vast amounts of heterogeneous biomedical data that are being continuously generated, such as in Electronic Health Records (EHRs). EHRs have been widely used in translational research to learn predictive models for discovery and disease management across varying patient cohorts. The very first step in such EHR-based applications often concerns patient cohort identification. Cohort identification involves the specification of a collection of eligibility criterion that needs to be transformed into a computable representation using the EHR’s semantic backbone (i.e., coding systems or ontologies) before queries can run against the EHR database. However, there are two critical barriers in performing effective cohort identification from large-scale EHRs. The first one is data (or semantic) heterogeneity, caused by a mixed utilization of coding systems. The second one is the quality of the semantic backbone or ontology hierarchy, which is essential for translating patient eligibility criteria to executable database queries. To address such challenges, this project will develop new methods for ontology matching and for ontology quality enhancement that directly impact data science practice in biomedicine, such as patient cohort identification. In addition, this project will incorporate the proposed computational aspects into data science-based courses to train next generation data scientists.This project consists of three research objectives. In Objective 1, the PI will develop new graph neural network (GNN)-based learning methods for matching biomedical ontologies by harnessing knowledge embedded in sources such as the Unified Medical Language System. This will address the heterogeneity issue and achieve semantic interoperability. In Objective 2, the PI will develop learning-based methods for detecting quality defects in subclass relations. This will address the quality issue and achieve continued enhancement of ontology hierarchies. In Objective 3, the PI will develop an ontology-based COVID-19 query engine for patient cohort identification, which is a real-world application of enhancing semantic interoperability for supporting data-driven COVID-19 research. For evaluation of the proposed methods, domain experts will be involved in validation of the resulted matching concepts and detected quality issues. The PI will communicate validated quality issues to the respective ontology owners for correction in subsequent ontology versions.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(11)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
A Query Engine for Self-controlled Case Series: with an application to COVID-19 EHR data
用于自我控制案例系列的查询引擎:适用于 COVID-19 EHR 数据
DOI:
--
发表时间:
2023
期刊:
AMIA Summits on Translational Science Proceedings
影响因子:
--
作者:
[Li, Xiaojin, Huang, Yan, Cui, Licong, Zhang, Guo-Qiang]
通讯作者:
Zhang, Guo-Qiang
DOI:
10.1093/bib/bbac122
发表时间:
2022-05-13
期刊:
Briefings in bioinformatics
影响因子:
9.5
作者:
[]
通讯作者:
A substring replacement approach for identifying missing IS-A relations in SNOMED CT
一种用于识别 SNOMED CT 中缺失 IS-A 关系的子串替换方法
DOI:
10.1109/bibm55620.2022.9995595
发表时间:
2023
期刊:
International Conference on Bioinformatics and Biomedicine
影响因子:
--
作者:
[Hao, Xubing, Abeysinghe, Rashmie, Shi, Jay, Cui, Licong]
通讯作者:
Cui, Licong
Identifying Missing IS-A Relations in Orphanet Rare Disease Ontology
识别孤儿罕见疾病本体中缺失的 IS-A 关系
DOI:
10.1109/bibm55620.2022.9995614
发表时间:
2022
期刊:
2022 IEEE International Conference on Bioinformatics and Biomedicine (BIBM
影响因子:
--
作者:
[Mohtashamian, Maryamsadat, Abeysinghe, Rashmie, Hao, Xubing, Cui, Licong]
通讯作者:
Cui, Licong
Automated Identification of Missing IS-A Relations in Human Phenotype Ontology
自动识别人类表型本体中缺失的 IS-A 关系
DOI:
--
发表时间:
2022
期刊:
AMIA Annual Symposium proceedings
影响因子:
--
作者:
[Mohtashamian, Maryamsadat, Hu, Ran, Abeysinghe, Rashmie, Hao, Xubing, Xu, Hua, Cui, Licong.]
通讯作者:
Cui, Licong.
共 7 条
III: Small: Methods for Auditing and Enhancing Completeness of Ontologies
-
批准号:1931134
-
项目类别:Standard Grant
-
资助金额:$30.69万
-
财政年份:2019
-
负责人:Licong Cui
-
依托单位:
III: Small: Methods for Auditing and Enhancing Completeness of Ontologies
-
批准号:1816805
-
项目类别:Standard Grant
-
资助金额:$30.95万
-
财政年份:2018
-
负责人:Licong Cui
-
依托单位:
CRII: III: A Scalable Framework for Debugging Large Biological Ontologies
-
批准号:1657306
-
项目类别:Standard Grant
-
资助金额:$15.1万
-
财政年份:2017
-
负责人:Licong Cui
-
依托单位:
海外基金