课题基金 / 基金详情

Use of NLP to Extract Risk Indicators for Immunologic Disease from the Text of EHRs (UNIITE)

Use of NLP to Extract Risk Indicators for Immunologic Disease from the Text of EHRs (UNIITE)
使用 NLP 从 EHR 文本中提取免疫疾病的风险指标 (UNIITE)
批准号:
10615338
负责人:
Nicholas L Rider
金额:
$20.94万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-02-18 至 2023-12-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
项目总结 卫生信息技术促进经济和临床卫生(HITECH)法案,使 采用电子健康记录(EHR)。在步调一致的情况下,医疗系统已经看到了影响 生物医学信息学技术,如用于挖掘推理的自然语言处理(NLP)和 对电子病历数据进行分类。将数字健康记录和可用的分析工具结合起来代表着一个机会 改善对患有罕见疾病的患者的护理,如原发免疫缺陷(PID) 最佳结果是建立在早期发现的基础上。然而,只有一小部分PID患者接受了 在遭受严重感染之前进行诊断。这强调了对改进诊断的新方法的需要 速度和驾驶了解的PID。目前,检测PID的障碍包括识别 不同的临床特征,区别于正常宿主的感染,以及普遍的缺乏 对疾病的认识。创建精确的分析方法,从EHR中挖掘并做出预测 数据代表了这些挑战的潜在解决方案。因此,我们的目标是开发一种自动控制系统 从EHR笔记中提取PID风险指标,以提高广泛的诊断和 提高对人类免疫性疾病的认识。我们的初步工作表明,结构化的电子病历 可以使用诸如问题列表元素和诊断代码之类的数据来开发概率框架 评估PID的风险,但它们是有限的,并且不能举例说明优化所需的全部概念 描述PID的特征。从文本中挖掘的数据元素可以耦合到当前可用的EHR结构化数据 改进的PID注释的本体。建立特定于PID的风险指标框架将使NLP PID检测的方法和对人类免疫功能障碍的更好理解。具体的 我们建议的目标如下:1)使用数据驱动的方法来识别和列举PID风险 来自EHR文本的指标。2.)开发自动提取关键PID风险指标的NLP方法 电子病历文本。我们的建议利用了从2000多名PID患者中捕获的非常大的EHR笔记文本语料库 在他们最终诊断之前,以及近5000名对照患者。挖掘该数据集,并协同 使用最先进的自然语言处理方法将为强大的、可互操作的文本挖掘奠定基础 用于PID风险检测的系统。我们希望这项工作将促进疾病的检测、特定疾病的特征 人类免疫缺陷,并允许结合临床、实验室和分子的额外推断 有关PID的信息。
英文摘要
PROJECT SUMMARY The Health Information Technology for Economic and Clinical Health (HITECH) Act, enabled widespread adoption of electronic health records (EHRs). In lockstep, healthcare systems have seen the impact of biomedical informatics techniques such as natural language processing (NLP) for mining inference and classifying EHR data. Combining digital health records and available analytical tools represents an opportunity to improve care for patients who suffer from rare disease such as primary immune deficiency (PID) where optimal outcomes are predicated upon early detection. However, only a fraction of PID patients receive a diagnosis before sustaining serious infections. This underscores a need for novel methods to improve diagnostic rates and for driving understanding about PID. At present, barriers to detecting PIDs include recognizing heterogeneous clinical features, distinguishing infections in PID from that of the normal host, and general lack of awareness about the diseases. Creating precise analytical methods to mine and make predictions from EHR data represents a potential solution to these challenges. As such our goals are to develop a system for automatic extraction of PID risk indicators from EHR notes for the purpose of improving widespread diagnosis and advancing knowledge about human immunologic disease. Our preliminary work suggests that structured EHR data such as problem list elements and diagnostic codes can be used to develop a probabilistic framework for assessing risk of PID but they are limited and do not exemplify the full range of concepts needed to optimally characterize PID. Data elements mined from text can couple to presently available EHR structured data and ontologies for improved annotation of PID. Building a framework of PID-specific risk indicators will enable NLP approaches for PID detection and improved understanding about human immune dysfunction. The Specific Aims for our proposal are as follows: 1.) To use a data-driven approach for identifying and enumerating PID risk indicators from EHR text. 2.) To develop NLP methods for automatically extracting key PID risk indicators from EHR text. Our proposal leverages a very large corpus of EHR note text captured from over 2000 PID patients prior to their ultimate diagnosis, as well as almost 5000 control patients. Mining this dataset and synergistically using state-of-the-art NLP methodologies will build the foundation of a potent and interoperable text mining system for PID risk detection. We expect this work to advance disease detection, characterization of specific human immune defects and allow for additional inference which combines clinical, laboratory, and molecular information about PID.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金