Leveraging Unlabeled and Pseudo Data for Clinical Information Extraction
Leveraging Unlabeled and Pseudo Data for Clinical Information Extraction
批准号:
9813134
负责人:
Ozlem Uzuner
金额:
$41.48万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-08-01 至 2022-07-31
关键词:
Accident and Emergency departmentAddressAdverse drug eventAffectClinicClinicalClinical DataCommunitiesComputer softwareDataData SetDevelopmentDiscipline of NursingElectronic Health RecordEngineeringEvaluationFrequenciesGoldGrowthHealthcareHospitalsInstitutionIsraelKnowledgeLabelLearningLinguisticsLocationMachine LearningMeasuresMedicalMedical centerMethodsModelingNamesNatural Language ProcessingNatureOutcomePatternPerformancePersonsPharmaceutical PreparationsPlant RootsProceduresPsychiatryPublicationsRecordsReportingResearchResourcesRouteSamplingSemanticsSigns and SymptomsSocial WorkStructureSupervisionSystemSystems DevelopmentTask PerformancesTelephoneTest ResultTestingTextThinnessTimeTrainingUniversitiesVariantVirginiaWashingtoncomputerizeddeep learningdosagefield studyimprovedlearning strategymedication administrationnovelopen sourceresponsesupervised learningtool
中文摘要
项目摘要/摘要
电子健康记录(EHR)包含重要的信息,这些信息可以使许多下游用户受益。
然而,这些信息大多是非结构化的叙事形式,无法通过计算机方法获取
它们依靠结构化表示来探索、检索和呈现信息。自然语言
处理(NLP)和信息提取(IE)为研究打开了这一信息宝库
不要这样。
在过去的几十年里,已经开发了许多IE系统。这些系统通常集中在一个
一次完成任务。此外,大多数只研究了特定类型的记录,例如,出院摘要,以及
根据来自单一机构的数据解决了他们的任务。由最先进的IE系统实现的性能
在此条件下开发的范围从44%的F-测量到99%的F-测量。这种观察到的变化可以
归因于任务的性质:一些目标实体,如日期,往往在数据中表现得更好
也更严格地坚持已知的表达方式,而不是用药的理由
它们在数据中相对稀疏,可以显示更广泛的语言多样性。然而,这可能不是唯一的
原因:使用的数据也可以解释性能差异。EHR的叙述有不同的风格,格式,
内容从一个科室到另一个科室,从一个医院到另一个医院。即使是相同的记录类型
两家不同的医院在叙事风格上可能非常不同,对IE提出了不同的挑战。
因此,要想了解IE的性能,需要研究针对多种记录类型的多项任务
来自多个机构。对如此大规模的IE系统进行评估的一个主要瓶颈是注释。
同样的瓶颈也限制了系统的开发。这项提议旨在解决这一瓶颈,为双方
评估和发展。它首先生成一个由多个记录类型组成的多机构语料库
五个机构。它研究了四种不同的IE任务,这些任务在临床记录中广泛代表IE,并可以为
IE整体领域:去辨别、临床概念提取、用药提取、药物不良事件
拔牙。在这些IE任务的背景下,提案然后提出了学习未标记的方法
或伪数据,这些数据可以帮助减少开发对带注释的数据的依赖。对这些方法进行评估
对于来自多个机构的多种类型的记录的性能和概括性。由于……
这些活动,该提案生成未识别的数据、注释、方法、软件和机器
学习模型,然后提供给研究社区。
英文摘要
Project Summary/Abstract
Electronic Health Records (EHRs) contain significant information that can benefit many downstream uses.
However, most of this information is in unstructured narrative form and is inaccessible to computerized methods
that rely on structured representations for exploring, retrieving, and presenting the information. Natural language
processing (NLP) and information extraction (IE) open this trove of information to studies that would otherwise
be without.
Over the past decades, many IE systems have been developed. These systems have typically focused on one
task at a time. In addition, most have studied only specific types of records, e.g., discharge summaries, and
addressed their task on data from a single institution. Performances achieved by the state-of-the-art IE systems
developed under these conditions ranged from 44% F-measure to 99% F-measure. This observed variation can
be attributed to the nature of the tasks: some target entities like dates tend to be better represented in the data
and also more rigidly stick to known patterns of expression as opposed to reasons for medication administration
which are relatively sparse in the data and can show wider linguistic diversity. However, this may not be the only
reason: the data used can also explain the performance variation. Narratives of EHRs vary in their style, format,
and content going from one department to another, from one hospital to another. Even the same record type in
two different hospitals can be very different in narrative style and pose different challenges for IE.
Understanding IE performance therefore requires studies of multiple tasks on multiple record types that come
from multiple institutions. One major bottleneck for evaluation of IE systems on such a large scale is annotation.
The same bottleneck also limits system development. This proposal aims to address this bottleneck for both
evaluation and development. It first generates a multi-institution corpus consisting of multiple record types from
five institutions. It studies four different IE tasks that broadly represent IE in clinical records and can inform the
field of IE as a whole: de-identification, clinical concept extraction, medication extraction, and adverse drug event
extraction. Within the context of these IE tasks, the proposal then puts forward methods that learn from unlabeled
or pseudo data that can help alleviate reliance on annotated data for development. It evaluates these methods
both for performance and generalizability on multiple types of records from multiple institutions. As a result of
these activities, this proposal generates de-identified data, annotations, methods, software, and machine
learning models which it then makes available to the research community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Joint learning methods for event and relation extraction from clinical narratives
-
批准号:10507223
-
项目类别:
-
资助金额:$42.49万
-
财政年份:2022
-
负责人:Ozlem Uzuner
-
依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
-
批准号:10670801
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
-
批准号:9759499
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
-
批准号:10393499
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2019
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing in Clinical Text
-
批准号:9597333
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2017
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8722031
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8400218
-
项目类别:
-
资助金额:$2.0万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
Challenges in Natural Language Processing for Clinical Narratives
-
批准号:8913773
-
项目类别:
-
资助金额:$1.98万
-
财政年份:2012
-
负责人:Ozlem Uzuner
-
依托单位:
海外基金