Use of NLP to Extract Risk Indicators for Immunologic Disease from the Text of EHRs (UNIITE)
Use of NLP to Extract Risk Indicators for Immunologic Disease from the Text of EHRs (UNIITE)
批准号:
10615338
负责人:
Nicholas L Rider
金额:
$20.94万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-02-18 至 2023-12-31
关键词:
AchievementAdoptionAlgorithmsAutomobile DrivingAwarenessBiologyCalibrationCaringCessation of lifeChildhoodClinicalClinical DataClinical ImmunologyCodeCollectionCoupledDataData ElementData SetDefectDetectionDiagnosisDiagnosticDiseaseEarly DiagnosisEconomicsElectronic Health RecordElementsFoundationsFutureGoalsHealthHealth PersonnelHealthcare SystemsHumanImmuneImmune System DiseasesImmunologicsImmunologyIndividualInfectionInternationalKnowledgeLaboratoriesLanguageLeadMachine LearningManualsMethodologyMethodsMiningModelingMolecularMorbidity - disease rateNatural Language ProcessingNatural Language Processing pipelineOntologyOutcomePatient CarePatient-Focused OutcomesPatientsPerformancePhenotypePositioning AttributeProcessRare DiseasesRiskRisk AssessmentSocietiesStatistical ModelsStructureSupervisionSystemTechniquesTestingTextTimeWorkadvanced diseaseanalytical methodanalytical toolbasebiomedical informaticsclinical practiceclinically relevantcohortcongenital immunodeficiencycostdigital healthdisabilityelectronic structurehealth information technologyhealth recordimprovedinnovationinsightinteroperabilitylarge datasetsmachine learning frameworkmachine learning modelmodel designmortalitynovelpredictive modelingscale upstructured datasupervised learningtext searchingtool
中文摘要
项目概要
《经济和临床健康健康信息技术 (HITECH) 法案》使广泛
采用电子健康记录(EHR)。医疗保健系统同步看到了影响
生物医学信息学技术,例如用于挖掘推理的自然语言处理(NLP)
对 EHR 数据进行分类。结合数字健康记录和可用的分析工具代表了一个机会
改善对患有原发性免疫缺陷 (PID) 等罕见疾病的患者的护理
最佳结果取决于早期发现。然而,只有一小部分 PID 患者接受了治疗
在严重感染之前进行诊断。这强调需要新的方法来改善诊断
率并促进对 PID 的理解。目前,检测 PID 的障碍包括识别
异质的临床特征,将 PID 感染与正常宿主的感染区分开来,并且普遍缺乏
提高人们对疾病的认识。创建精确的分析方法来挖掘 EHR 并做出预测
数据代表了这些挑战的潜在解决方案。因此,我们的目标是开发一个自动系统
从 EHR 记录中提取 PID 风险指标,以改善广泛的诊断和
增进有关人类免疫疾病的知识。我们的初步工作表明,结构化 EHR
问题列表元素和诊断代码等数据可用于开发概率框架
评估 PID 风险,但它们是有限的,并且没有体现最佳化所需的全部概念
表征PID。从文本中挖掘的数据元素可以与当前可用的 EHR 结构化数据相结合,
用于改进 PID 注释的本体。构建 PID 特定风险指标框架将使 NLP 成为可能
PID 检测方法并提高对人类免疫功能障碍的了解。具体
我们建议的目标如下: 1.) 使用数据驱动的方法来识别和枚举 PID 风险
EHR 文本中的指标。 2.) 开发 NLP 方法,自动提取关键 PID 风险指标
电子病历文本。我们的提案利用了从 2000 多名 PID 患者中捕获的大量 EHR 注释文本语料库
在最终诊断之前,以及近 5000 名对照患者。协同挖掘该数据集
使用最先进的 NLP 方法将为有效且可互操作的文本挖掘奠定基础
PID风险检测系统。我们期望这项工作能够推进疾病检测、特定疾病的表征
人类免疫缺陷,并允许结合临床、实验室和分子的额外推断
有关 PID 的信息。
英文摘要
PROJECT SUMMARY
The Health Information Technology for Economic and Clinical Health (HITECH) Act, enabled widespread
adoption of electronic health records (EHRs). In lockstep, healthcare systems have seen the impact
of biomedical informatics techniques such as natural language processing (NLP) for mining inference and
classifying EHR data. Combining digital health records and available analytical tools represents an opportunity
to improve care for patients who suffer from rare disease such as primary immune deficiency (PID) where
optimal outcomes are predicated upon early detection. However, only a fraction of PID patients receive a
diagnosis before sustaining serious infections. This underscores a need for novel methods to improve diagnostic
rates and for driving understanding about PID. At present, barriers to detecting PIDs include recognizing
heterogeneous clinical features, distinguishing infections in PID from that of the normal host, and general lack
of awareness about the diseases. Creating precise analytical methods to mine and make predictions from EHR
data represents a potential solution to these challenges. As such our goals are to develop a system for automatic
extraction of PID risk indicators from EHR notes for the purpose of improving widespread diagnosis and
advancing knowledge about human immunologic disease. Our preliminary work suggests that structured EHR
data such as problem list elements and diagnostic codes can be used to develop a probabilistic framework for
assessing risk of PID but they are limited and do not exemplify the full range of concepts needed to optimally
characterize PID. Data elements mined from text can couple to presently available EHR structured data and
ontologies for improved annotation of PID. Building a framework of PID-specific risk indicators will enable NLP
approaches for PID detection and improved understanding about human immune dysfunction. The Specific
Aims for our proposal are as follows: 1.) To use a data-driven approach for identifying and enumerating PID risk
indicators from EHR text. 2.) To develop NLP methods for automatically extracting key PID risk indicators from
EHR text. Our proposal leverages a very large corpus of EHR note text captured from over 2000 PID patients
prior to their ultimate diagnosis, as well as almost 5000 control patients. Mining this dataset and synergistically
using state-of-the-art NLP methodologies will build the foundation of a potent and interoperable text mining
system for PID risk detection. We expect this work to advance disease detection, characterization of specific
human immune defects and allow for additional inference which combines clinical, laboratory, and molecular
information about PID.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金