Structured deep embedding model to generate composite clinical indices from electronic health records for early detection of pancreatic cancer.

Structured deep embedding model to generate composite clinical indices from electronic health records for early detection of pancreatic cancer.
复制标题

DOI:
10.1016/j.patter.2022.100636
复制
发表时间:
2023-01-13
期刊:
影响因子:
6.5
通讯作者:
Tatonetti, Nicholas P.
Tatonetti, Nicholas P.
中科院分区:
其他
文献类型:
--
作者:
Park, Jiheum;Artin, Michael G.;Lee, Kate E.;May, Benjamin L.;Park, Michael;Hur, Chin;Tatonetti, Nicholas P.

文献摘要

参考文献

被引文献

相似文献

电子健康记录(EHR)数据的高维性、复杂性和不规则性给简化和全面的健康评估带来了重大挑战,阻碍了临床医生有效提取可操作的见解。如果我们能为人类决策者提供一组简化的可解释的综合指数(即,将关于相关测量的组的信息组合成单个代表值),这将促进有效的临床决策。在这项研究中,我们建立了一个结构化的深度嵌入模型,旨在通过对领域专家确定的相关测量进行分组来降低输入变量的维度(例如,临床医生)。我们的研究结果表明,代表肝功能的综合指标可能始终是胰腺癌(PC)早期检测的最重要因素。我们提出我们的模型作为利用深度学习从EHR开发综合指数的基础,用于预测健康结果,包括但不限于各种癌症,并具有临床意义的解释。我们的研究表明,基于深度学习的方法用于生成复合指数(CI)该方法结合了领域知识融合的方法相关信息分组的策略决定了CI的可解释性领域知识知情的CI可以促进临床决策患者病历数量的爆炸性增长导致医疗保健提供者的信息过载,并且当前存在与利用该信息来提高临床决策的质量相关联的许多重大挑战。在这项工作中,我们的目标是开发新的简化表示的病人状态,这是预测和解释的医生。这些患者状态表示通过深度学习架构进行聚合,该架构利用领域知识将每位患者可用的大量临床变量分组为一组简化的复合指数,同时保留解释模型如何得出最终预测的能力。我们预计,我们的领域知识融合的方法将提供一个基础,产生新的可解释的高层次的综合指数,减少“黑箱”的关注模型的有效性,从而提高临床采用决策。这项研究提出了一种生成式深度学习模型的概念证明,该模型可以将大量的EHR变量提取为一组易于处理的综合指数,这些指数以临床上有意义和人类可解释的形式量化癌症风险。以胰腺癌早期检测为具体目标,我们从EHR的206个临床时间序列变量中产生了5个器官特异性复合指数,代表肝功能的复合指数始终被证明是最重要的预测因子。
The high-dimensionality, complexity, and irregularity of electronic health records (EHR) data create significant challenges for both simplified and comprehensive health assessments, prohibiting an efficient extraction of actionable insights by clinicians. If we can provide human decision-makers with a simplified set of interpretable composite indices (i.e., combining information about groups of related measures into single representative values), it will facilitate effective clinical decision-making. In this study, we built a structured deep embedding model aimed at reducing the dimensionality of the input variables by grouping related measurements as determined by domain experts (e.g., clinicians). Our results suggest that composite indices representing liver function may consistently be the most important factor in the early detection of pancreatic cancer (PC). We propose our model as a basis for leveraging deep learning toward developing composite indices from EHR for predicting health outcomes, including but not limited to various cancers, with clinically meaningful interpretations. Our study shows deep-learning-based approach for generating composite indices (CIs) The approach incorporates the method for domain-knowledge fusion A strategy for grouping relevant information determines CI interpretability The domain-knowledge-informed CI can facilitate clinical decision-making The explosive growth in the volume of patient medical records has resulted in an overload of information for healthcare providers, and there are currently numerous significant challenges associated with leveraging this information to improve the quality of clinical decision-making. In this work, we aim to develop new simplified representations of patient states that are both predictive and interpretable to physicians. These patient state representations are aggregated via deep-learning architectures that leverage domain knowledge to group the large number of clinical variables available per patient into a simplified set of composite indices while preserving the ability to explain how the model arrived at the final prediction. We anticipate that our methods for domain-knowledge fusion will provide a basis for producing new interpretable high-level composite indices that reduce “black box” concerns regarding model validity and therefore improve clinical adoption into decision-making. This study presents the proof of concept for a generative deep-learning model that can distill a massive volume of EHR variables into a tractable set of composite indices that quantify the risk of cancer in a clinically meaningful and human interpretable form. With the specific aim of the early detection of pancreatic cancer, we generated five organ-specific composite indices out of 206 clinical time-series variables from EHR, and the composite index representing liver function was consistently shown to be the most important predictor.
DOI: 10.1016/j.semradonc.2019.05.010
发表时间: 2019-10-01
影响因子: 3.5
作者:
Kim, Ellen;Rubinstein, Samuel M.;Warner, Jeremy L.
通讯作者: Warner, Jeremy L.
DOI: 10.1097/mpa.0000000000001882
发表时间: 2021-08-01
期刊: Pancreas
影响因子: 2.9
作者:
Kenner BJ;Abrams ND;Chari ST;Field BF;Goldberg AE;Hoos WA;Klimstra DS;Rothschild LJ;Srivastava S;Young MR;Go VLW
通讯作者: Go VLW
DOI: 10.1093/ajcp/aqw064
发表时间: 2016-06-01
影响因子: 3.5
作者:
Luo, Yuan;Szolovits, Peter;Baron, Jason M.
通讯作者: Baron, Jason M.
DOI: 10.3109/0886022x.2012.718951
发表时间: 2012-01-01
期刊: RENAL FAILURE
影响因子: 3
作者:
Golay, Vishal;Roychowdhary, Arpita
通讯作者: Roychowdhary, Arpita
Med-BERT:在大规模结构化电子健康记录上进行预训练的情境化嵌入,用于疾病预测。
DOI: 10.1038/s41746-021-00455-y
发表时间: 2021-05-20
影响因子: 15.2
作者:
Rasmy L;Xiang Y;Xie Z;Tao C;Zhi D
通讯作者: Zhi D