课题基金 / 基金详情

项目摘要

项目成果

Ozlem Uzuner的其他基金

相似基金

相关文献

中文摘要
翻译
项目摘要 电子健康记录(EHR),详细说明患者状态和临床护理的各个方面,可以极大地促进 质量改进和监督举措,以及革命性的临床研究。非结构化的 电子病历中的临床叙述记录关键信息,包括医疗问题、治疗和 诊断性测试以及护理和结果的理由。自然语言处理(NLP)和 信息提取(IE)系统的目标是从临床叙述中识别这些关键信息。 这些系统提取临床概念,如医疗问题、治疗和测试,确定 这些概念的属性,以明确它们的存在/不存在以及患者的其他细节;并确定 根据预定义的关系,这些概念之间的相互作用。大多数临床NLP系统 处理这些信息的提取是基于管道的:即,临床概念的提取先于 确定它们的属性和确定临床概念之间的关系。在制作过程中 这些系统前景看好,但存在两大局限性:(1)当面临数据不平衡时,它们 在数据中发现的更普遍的观察类别上表现最好,而在不太普遍的观察类别上表现最差 一,以及(2)它们允许组件之间的错误级联。这两个限制还可以 相互复合。因此,NLP系统提取的信息可能是不完整和粗糙的- 细粒度的,无法支持需要更细粒度的患者病情图像的临床应用程序。 在这个项目中,我们建议解决临床信息提取任务中的这些限制,该任务旨在 使用新颖、细粒度的分层架构,更全面地了解患者的病情 临床上突出的事件及其关系。我们将临床显著事件定义为医疗问题, 在病人护理过程中记录的治疗和测试。我们在一个包含以下内容的帧中捕获每个事件 触发器和一组细粒度属性。我们在事件之上建立事件-事件关系。致信地址 针对数据不平衡问题,我们提出了(I)一种新的主动学习框架,用于指导人工标注工作 走向多样化和信息量丰富的样本,以提高对不太普遍的属性的自动识别,并 关系。为了解决级联错误,我们提出了一种新的支持多任务的联合学习系统 相互告知,以便在所有任务中获得更好的性能。我们评估我们在多种音符类型上的工作 来自多个机构。预期成果包括:(1)一个全面的、不同种类的黄金标准 从多个机构为临床显著事件和关系创建的数据集,(2)NLP方法 生成最先进的结果来提取事件和关系,以及(3)记录我们的 调查结果。批注指南和模式、黄金标准批注以及NLP模型和工具 在项目期间创建的数据将与研究社区共享。
英文摘要
Project Summary Electronic health records (EHRs), detailing patient status and all aspects of clinical care, can greatly facilitate quality improvement and surveillance initiatives as well as revolutionize clinical research. The unstructured clinical narratives in EHRs document critical information, including medical problems, treatments, and diagnostic tests as well as the rationale for care and outcomes. Natural Language Processing (NLP) and Information Extraction (IE) systems target the identification of such critical information from clinical narratives. These systems extract clinical concepts such as medical problems, treatments, and tests, determine the attributes of these concepts to get clarity on their presence/absence and other details in a patient; and identify the interactions of these concepts with each other in terms of predefined relations. Most clinical NLP systems that tackle the extraction of this information are pipeline based: i.e., extraction of clinical concepts precedes the determination of their attributes and the determination of relations between clinical concepts. While producing promising results, these systems suffer from two major limitations: (1) when faced with data imbalance, they perform best on the more prevalent classes of observations found in the data and suffer on the less prevalent ones, and (2) they allow errors to cascade between the components. These two limitations can also compound each other. As a result, the information extracted by NLP systems can be incomplete and coarse- grained, unable to support clinical applications that require a more fine-grained picture of the patient condition. In this project, we propose to address these limitations on a clinical information extraction task that aims to capture a more complete picture of the patient condition with a novel, fine-grained, hierarchical schema for clinically-salient events and their relations. We define clinically-salient events as medical problems, treatments, and tests that are documented during patient care. We capture each event in a frame that consists of a trigger and a set of fine-grained attributes. We build event–event relations on top of events. To address data imbalance, we propose (i) a novel active learning framework that guides manual annotation efforts towards diverse and informative samples that can boost automated recognition of less prevalent attributes and relations. To address cascading errors, we propose (ii) a novel joint learning system that enables multiple tasks to inform each other for better performance across all tasks. We evaluate our work on multiple note types from multiple institutions. Expected outcomes include (1) a comprehensive heterogeneous gold-standard dataset created from multiple institutions for clinically-salient events and relations, (2) NLP methods that generate state-of-the-art results in extraction of events and relations, and (3) publications that document our findings. The annotation guidelines and schema, the gold-standard annotations, and the NLP models and tools created during the project will be shared with the research community.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1038/s41597-022-01521-0
发表时间: 2022-08-11
期刊: SCIENTIFIC DATA
影响因子: 9.8
作者: [Dobbins, Nicholas J., Mullen, Tony, Uzuner, Ozlem, Yetisgen, Meliha]
通讯作者: Yetisgen, Meliha
MT-clinical BERT: scaling clinical information extraction with multitask learning.
MT-clinical BERT:通过多任务学习扩展临床信息提取。
DOI: 10.1093/jamia/ocab126
发表时间: 2021
期刊: Journal of the American Medical Informatics Association : JAMIA
影响因子: --
作者: [Mulyar,Andriy, Uzuner,Ozlem, McInnes,Bridget]
通讯作者: McInnes,Bridget
DOI: 10.1016/j.jbi.2020.103552
发表时间: 2020-10
期刊: Journal of biomedical informatics
影响因子: 4.5
作者: [Sutphin C, Lee K, Yepes AJ, Uzuner Ö, McInnes BT]
通讯作者: McInnes BT
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
  • 批准号:
    10670801
  • 项目类别:
  • 资助金额:
    $2.0万
  • 财政年份:
    2019
  • 负责人:
    Ozlem Uzuner
  • 依托单位:
Leveraging Unlabeled and Pseudo Data for Clinical Information Extraction
  • 批准号:
    9813134
  • 项目类别:
  • 资助金额:
    $41.48万
  • 财政年份:
    2019
  • 负责人:
    Ozlem Uzuner
  • 依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
  • 批准号:
    9759499
  • 项目类别:
  • 资助金额:
    $2.0万
  • 财政年份:
    2019
  • 负责人:
    Ozlem Uzuner
  • 依托单位:
National NLP Clinical Challenges (n2c2): Challenges in Natural Language Processing for Clinical Narratives
  • 批准号:
    10393499
  • 项目类别:
  • 资助金额:
    $2.0万
  • 财政年份:
    2019
  • 负责人:
    Ozlem Uzuner
  • 依托单位:
海外基金