Using natural language processing and machine learning to identify patients at high risk of potentially preventable carbapenem-resistant Enterobacterales (CRE) infection
Using natural language processing and machine learning to identify patients at high risk of potentially preventable carbapenem-resistant Enterobacterales (CRE) infection
批准号:
10705635
负责人:
Katherine Elizabeth Goodman
金额:
$12.36万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-09-30 至 2026-09-29
中文摘要
项目摘要
耐碳青霉烯类肠杆菌(CRE)每年在美国住院患者中导致超过1.3万人感染,
死亡率可以超过50%。在医院里,无症状的殖民正在成为一个关键
预防CRE感染的目标:从流行病学上讲,及早识别移居的患者可以减少
在医院内传播。临床上,被殖民的住院患者面临的风险要高得多--但可能是可以改变的
-Cre感染的风险。然而,由于诊断的局限性,大规模Cre定植筛查
对大多数美国医院来说仍然不切实际。预测模型提供了可供选择的策略
有很高的定居风险和随后感染风险的患者。然而,模型面临两种方法论
障碍,限制了其更广泛的用途:(1)强烈的殖民风险因素被“锁定”在电子健康记录中
(EHR)自由文本,除非手动查看记录,否则无法用于建模;以及(2)由于
对于CRE患病率,大多数统计模型只能评估有限数量的候选变量。
我们建议利用最先进的机器学习和自然语言处理(NLP)技术来
提高对CRE定植和感染患者的识别能力。我们将把这些方法应用于来自
>;在马里兰大学和约翰霍普金斯医院对2.1万名患者进行了CRE筛查。在目标1中,
我们将在入院历史上构建并验证NLP算法,以检测入院前暴露于
很强的殖民风险因素,但在结构化的EHR数据字段中捕获较少。NLP是一个尖端的
“解锁”这些类型的非结构化数据的计算技术。我们还将使用文本挖掘
确定潜在的新的或当地的CRE风险因素的方法。在目标2和目标3中,我们将构建和验证
来自NLP派生变量和其他EHR数据的模型,以预测入院时的殖民(目标2)和
进展到感染(目标3)使用机器学习算法,在高维数据上表现出色。
综上所述,这项工作将帮助医院识别CRE定植和感染的高风险患者
早期,当有害的患者结果仍然可以预防的时候。因为NLP是自动化的、成功的模型
可以输出到其他医院并集成到EHR中;这项工作产生的所有算法将
免费提供的。这将是部署NLP进行细菌携带筛查的第一项研究,也是规模最大的
美国研究跟踪CRE殖民住院患者的感染情况。作为一名受过博士培训的流行病学家,拥有CRE和
有机器学习背景,并曾在FDA执业法律,我对跨学科很感兴趣,严谨
减少住院患者抗生素耐药死亡的途径和政策。在短期内,
职业发展奖的支持将使我能够使用复杂的计算来积累经验
基于电子病历的信息提取和预测建模方法。从长远来看,我获得的技能
将使我成为利用新的医疗信息技术工具来设计和测试新的
应对美国医院中抗生素耐药性和其他新出现病原体的威胁的策略。
英文摘要
Project Summary
Carbapenem-resistant Enterobacterales (CRE) cause more than 13,000 infections in U.S. inpatients annually,
with mortality rates that can exceed 50%. In hospitals, asymptomatic colonization is emerging as a critical
target for CRE infection prevention: epidemiologically, early identification of colonized patients can reduce
intra-hospital spread. And clinically, colonized inpatients face significantly higher — but potentially modifiable
— risks of CRE infection. Due to diagnostic limitations, however, widescale CRE colonization screening
remains impractical for most U.S. hospitals. Prediction models offer alternative strategies for identifying
patients at high risk of colonization and of subsequent infection. However, models face two methodological
obstacles, limiting their wider utility: (1) strong colonization risk factors are “locked” in electronic health record
(EHR) free-text that is unavailable for model-building unless records are reviewed manually; and (2) due to low
CRE prevalence, most statistical models can only evaluate limited numbers of candidate variables.
We propose to exploit state-of-the-art machine learning and natural language processing (NLP) techniques to
improve identification of CRE-colonized and infected patients. We will apply these methods to EHRs from
>21,000 patients screened for CRE at The University of Maryland and The Johns Hopkins hospitals. In Aim 1,
we will build and validate NLP algorithms on admission histories to detect pre-admission exposures that are
strong colonization risk factors but poorly captured in structured EHR data fields. NLP is a cutting-edge
computational technique for “unlocking” these types of unstructured data. We will also use text-mining
approaches to identify potential new or local CRE risk factors. In Aims 2 and 3, we will build and validate
models from NLP-derived variables and other EHR data to predict colonization at admission (Aim 2) and
progression to infection (Aim 3) using machine learning algorithms that excel on high-dimensional data.
Taken together, this work will help hospitals identify patients at high risk of CRE colonization and infection
early, when deleterious patient outcomes are still preventable. Because NLP is automated, successful models
could be exported to other hospitals and integrated into EHRs; all algorithms resulting from this work will be
made freely available. This will be the first study to deploy NLP for bacterial carriage screening and the largest
U.S. study to follow CRE-colonized inpatients for infection. As a PhD-trained epidemiologist with a CRE and
machine learning background, and who previously practiced FDA law, I am drawn to interdisciplinary, rigorous
approaches and policies for reducing the toll of antibiotic resistance in hospitalized patients. In the short-term,
Career Development Award support would allow me to build experience using sophisticated computational
approaches for EHR-based information extraction and predictive modeling. In the long-term, the skills I acquire
would position me as a leader at leveraging novel health information technology tools to design and test new
strategies for responding to the threat of antibiotic resistance and other emerging pathogens in U.S. hospitals.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
Using natural language processing and machine learning to identify patients at high risk of potentially preventable carbapenem-resistant Enterobacterales (CRE) infection
-
批准号:10370475
-
项目类别:
-
资助金额:$13.03万
-
财政年份:2021
-
负责人:Katherine Elizabeth Goodman
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Natural超对称中的希格斯物理与暗物质研究
-
批准号:11775039
-
项目类别:面上项目
-
资助金额:52.0万元
-
批准年份:2017
-
负责人:郑思波
-
依托单位:
Natural超对称在LHC上的现象学研究
-
批准号:11405015
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2014
-
负责人:郑思波
-
依托单位:
双硅化合物反应及天然产物合成应用研究
-
批准号:21172150
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2011
-
负责人:宋振雷
-
依托单位:
受体编辑在天然自身反应性B细胞发育耐受中的作用和机制研究
-
批准号:30901336
-
项目类别:青年科学基金项目
-
资助金额:18.0万元
-
批准年份:2009
-
负责人:邢影
-
依托单位:
海洋天然产物Amphidinolide G和H全合成研究
-
批准号:20772148
-
项目类别:面上项目
-
资助金额:30.0万元
-
批准年份:2007
-
负责人:赵刚
-
依托单位:
抗肾小球基底膜抗体的免疫学特性在疾病发生和发展中的作用
-
批准号:30700752
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2007
-
负责人:崔昭
-
依托单位: