Joint learning methods for event and relation extraction from clinical narratives
Joint learning methods for event and relation extraction from clinical narratives
批准号:
10507223
负责人:
Ozlem Uzuner
金额:
$42.49万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-10 至 2025-07-31
关键词:
Active LearningAddressAnaphylaxisAutomated AnnotationCaringClinicalClinical ResearchCodeCommunitiesComplementDataData SetDiagnosisDiagnostic testsDiseaseElectronic Health RecordEvaluationEventExhibitsGoldGrainGuidelinesHospitalsInformation RetrievalInfusion proceduresInstitutionIntensive Care UnitsIsraelJointsLearningManualsMedicalMedical centerMethodsModelingNatural Language ProcessingOutcomePatient CarePatientsPerformancePublicationsResearchRouteSamplingSemanticsStructureSystemTestingTextTrainingTriageUniversitiesVitamin KWashingtonWorkbasecare outcomesclinical applicationclinical careclinically significantdeep learningeffective care managementlearning strategymodel developmentnoveloutcome predictionpatient populationtool
中文摘要
项目总结
英文摘要
Project Summary
Electronic health records (EHRs), detailing patient status and all aspects of clinical care, can greatly facilitate
quality improvement and surveillance initiatives as well as revolutionize clinical research. The unstructured
clinical narratives in EHRs document critical information, including medical problems, treatments, and
diagnostic tests as well as the rationale for care and outcomes. Natural Language Processing (NLP) and
Information Extraction (IE) systems target the identification of such critical information from clinical narratives.
These systems extract clinical concepts such as medical problems, treatments, and tests, determine the
attributes of these concepts to get clarity on their presence/absence and other details in a patient; and identify
the interactions of these concepts with each other in terms of predefined relations. Most clinical NLP systems
that tackle the extraction of this information are pipeline based: i.e., extraction of clinical concepts precedes the
determination of their attributes and the determination of relations between clinical concepts. While producing
promising results, these systems suffer from two major limitations: (1) when faced with data imbalance, they
perform best on the more prevalent classes of observations found in the data and suffer on the less prevalent
ones, and (2) they allow errors to cascade between the components. These two limitations can also
compound each other. As a result, the information extracted by NLP systems can be incomplete and coarse-
grained, unable to support clinical applications that require a more fine-grained picture of the patient condition.
In this project, we propose to address these limitations on a clinical information extraction task that aims to
capture a more complete picture of the patient condition with a novel, fine-grained, hierarchical schema for
clinically-salient events and their relations. We define clinically-salient events as medical problems,
treatments, and tests that are documented during patient care. We capture each event in a frame that consists
of a trigger and a set of fine-grained attributes. We build event–event relations on top of events. To address
data imbalance, we propose (i) a novel active learning framework that guides manual annotation efforts
towards diverse and informative samples that can boost automated recognition of less prevalent attributes and
relations. To address cascading errors, we propose (ii) a novel joint learning system that enables multiple tasks
to inform each other for better performance across all tasks. We evaluate our work on multiple note types
from multiple institutions. Expected outcomes include (1) a comprehensive heterogeneous gold-standard
dataset created from multiple institutions for clinically-salient events and relations, (2) NLP methods that
generate state-of-the-art results in extraction of events and relations, and (3) publications that document our
findings. The annotation guidelines and schema, the gold-standard annotations, and the NLP models and tools
created during the project will be shared with the research community.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1038/s41597-022-01521-0
发表时间:
2022-08-11
期刊:
SCIENTIFIC DATA
影响因子:
9.8
作者:
[Dobbins, Nicholas J., Mullen, Tony, Uzuner, Ozlem, Yetisgen, Meliha]
通讯作者:
Yetisgen, Meliha
MT-clinical BERT: scaling clinical information extraction with multitask learning.
MT-clinical BERT:通过多任务学习扩展临床信息提取。
DOI:
10.1093/jamia/ocab126
发表时间:
2021
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
作者:
[Mulyar,Andriy, Uzuner,Ozlem, McInnes,Bridget]
通讯作者:
McInnes,Bridget
DOI:
10.1016/j.jbi.2020.103552
发表时间:
2020-10
期刊:
Journal of biomedical informatics
影响因子:
4.5
作者:
[Sutphin C, Lee K, Yepes AJ, Uzuner Ö, McInnes BT]
通讯作者:
McInnes BT
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
-
批准号:10670801
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
Leveraging Unlabeled and Pseudo Data for Clinical Information Extraction
-
批准号:9813134
-
项目类别:
-
资助金额:$41.48万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
-
批准号:9759499
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
-
批准号:10393499
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing in Clinical Text
-
批准号:9597333
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2017
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8722031
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8913773
-
项目类别:
-
资助金额:$1.98万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8400218
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
海外基金