Medication based machine learning to identify subpopulations of pediatric hemodialysis patients in an electronic health record database.

Medication based machine learning to identify subpopulations of pediatric hemodialysis patients in an electronic health record database.
复制标题

基于药物的机器学习,用于识别电子健康记录数据库中儿科血液透析患者的亚群。

DOI:
10.1016/j.imu.2022.101104
复制
发表时间:
2022
影响因子:
--
通讯作者:
Brewer,SimonC
Brewer,SimonC
中科院分区:
--
文献类型:
--
作者:
McKnite,AutumnM;Job,KathleenM;Nelson,Raoul;Sherwin,CatherineMT;Watt,KevinM;Brewer,SimonC

文献摘要

相似文献

电子健康记录(EHR)已经产生了庞大而复杂的医疗信息数据库,这些数据库有可能成为临床研究的强大工具。然而,不同机构在编码系统以及EHR内信息的细节和准确性方面存在差异。这使得确定患者亚群具有挑战性,并限制了多机构数据库的广泛使用。在这项研究中,我们利用机器学习来识别接受肾脏替代治疗的住院儿童患者的用药模式,并创建了一个成功区分间歇性(IHD)和持续性肾脏替代疗法(CRRT)血液透析患者的预测模型。我们训练了六种机器学习算法(逻辑回归、朴素贝叶斯、k近邻、支持向量机、随机森林和梯度增强树),使用来自多中心数据库的患者记录(n=533)和处方药物成分(n=288)作为特征来区分两种血液透析类型。预测技能使用5次交叉验证进行评估,算法的性能范围从0.7平衡精度(逻辑回归)到0.86(随机森林)。使用独立的单中心数据集对两个性能最好的模型进行了进一步测试,获得了84-87%的均衡准确率。该模型克服了大型数据库固有的问题,并将允许我们利用和结合历史记录,显著增加IHD和CRRT人群中的种群规模和多样性,用于未来的临床研究。我们的工作证明了单独使用药物来准确区分大数据集中的患者亚群的效用,允许在不同的编码系统之间传输代码。这一框架有可能被用来区分没有歧视性ICD编码的其他患者亚群,从而允许更详细的见解和新的研究路线。
Electronic health records (EHRs) have given rise to large and complex databases of medical information that have the potential to become powerful tools for clinical research. However, differences in coding systems and the detail and accuracy of the information within EHRs can vary across institutions. This makes it challenging to identify subpopulations of patients and limits the widespread use of multi-institutional databases. In this study, we leveraged machine learning to identify patterns in medication usage among hospitalized pediatric patients receiving renal replacement therapy and created a predictive model that successfully differentiated between intermittent (iHD) and continuous renal replacement therapy (CRRT) hemodialysis patients. We trained six machine learning algorithms (logistical regression, Naïve Bayes,k-nearest neighbor, support vector machine, random forest, and gradient boosted trees) using patient records from a multi-center database (n= 533) and prescribed medication ingredients (n= 228) as features to discriminate between the two hemodialysis types. Predictive skill was assessed using a 5-fold cross-validation, and the algorithms showed a range of performance from 0.7 balanced accuracy (logistical regression) to 0.86 (random forest). The two best performing models were further tested using an independent single-center dataset and achieved 84–87% balanced accuracy. This model overcomes issues inherent within large databases and will allow us to utilize and combine historical records, significantly increasing population size and diversity within both iHD and CRRT populations for future clinical studies. Our work demonstrates the utility of using medications alone to accurately differentiate subpopulations of patients in large datasets, allowing codes to be transferred between different coding systems. This framework has the potential to be used to distinguish other subpopulations of patients where discriminatory ICD codes are not available, permitting more detailed insights and new lines of research.