Predicting seizure recurrence after an initial seizure-like episode from routine clinical notes using large language models: a retrospective cohort study.

Predicting seizure recurrence after an initial seizure-like episode from routine clinical notes using large language models: a retrospective cohort study.
复制标题

DOI:
10.1016/s2589-7500(23)00179-6
复制
发表时间:
2023-12
影响因子:
30.8
通讯作者:
Kohane, Isaac
Kohane, Isaac
中科院分区:
医学1区
文献类型:
--
作者:
Beaulieu-Jones, Brett K.;Villamar, Mauricio F.;Scordis, Phil;Bartmann, Ana Paula;Ali, Waqar;Wissel, Benjamin;Alsentzer, Emily;de Jong, Johann;Patra, Arijit;Kohane, Isaac

文献摘要

参考文献

相似文献

评估和管理儿童首次癫痫样发作可能很困难,因为这些发作并不总是直接观察到的,可能是癫痫发作或其他情况(癫痫样发作)。我们的目的是评估使用真实世界数据的机器学习模型是否可以预测最初的癫痫样事件后癫痫的复发。这项回溯性队列研究比较了2010年1月1日和2020年1月1日在两个不同的数据集上训练和评估的模型:波士顿儿童医院的电子医疗记录(EMR)和IBM MarketScan研究数据库中未识别的患者级别的行政索赔数据。研究人群包括根据国际疾病分类、临床修改(ICD-CM)代码在21岁之前首次诊断为癫痫或惊厥的患者。我们比较了使用结构化数据(Logistic回归和XGBoost)的基于机器学习的预测建模与使用大型语言模型进行自然语言处理的新兴技术。波士顿儿童医院的主要队列包括14021名符合纳入标准的患者和最初的癫痫样事件,而比较队列包括IBM MarketScan研究数据库中的15062名患者。根据专家得出的复合定义,波士顿儿童医院57%的患者和IBM MarketScan的63%的患者癫痫复发。对于被排除在研究之外的患者,大语言模型加上额外的特定领域和特定位置的预训练(F1-Score 0.826[95%CI 0·817-0.835],AUC 0.897[95%CI 0·875-0·913])表现最好。所有大型语言模型,包括没有额外预训练的基础模型(F1-Score 0·739[95%CI 0·738-0·741],AUROC 0·846[95%CI 0·826-0·861])都优于用结构化数据训练的模型。仅在结构化数据下,XGBoost优于Logistic回归和用波士顿儿童医院EMR训练的XGBoost模型(Logistic回归:F1-分数0·650[95%CI 0·643-0·657],AUC 0·694[95%CI 0·685-0·705],XGBoost:F1-分数0·679[0·676-0·683],AUC 0·725[0·717-0·734])与在IBM MarketScan数据库上训练的模型(Logistic回归:F1-分数0·596[0·590-0·601])相似,AUC 0·670[0·664-0·675],XGBoost:F1-Score 0·678[0·668-0·687],AUC 0·710[0·703-0·714]。医生关于最初癫痫样事件的临床记录包括用于预测癫痫复发的大量信号,以及额外的特定领域和特定位置的预培训可以显著提高临床大型语言模型的性能,即使是对专门的队列也是如此。美国国家神经疾病和中风研究所(美国国立卫生研究院)。
The evaluation and management of first-time seizure-like events in children can be difficult because these episodes are not always directly observed and might be epileptic seizures or other conditions (seizure mimics). We aimed to evaluate whether machine learning models using real-world data could predict seizure recurrence after an initial seizure-like event. This retrospective cohort study compared models trained and evaluated on two separate datasets between Jan 1, 2010, and Jan 1, 2020: electronic medical records (EMRs) at Boston Children’s Hospital and de-identified, patient-level, administrative claims data from the IBM MarketScan research database. The study population comprised patients with an initial diagnosis of either epilepsy or convulsions before the age of 21 years, based on International Classification of Diseases, Clinical Modification (ICD-CM) codes. We compared machine learning-based predictive modelling using structured data (logistic regression and XGBoost) with emerging techniques in natural language processing by use of large language models. The primary cohort comprised 14 021 patients at Boston Children’s Hospital matching inclusion criteria with an initial seizure-like event and the comparison cohort comprised 15 062 patients within the IBM MarketScan research database. Seizure recurrence based on a composite expert-derived definition occurred in 57% of patients at Boston Children’s Hospital and 63% of patients within IBM MarketScan. Large language models with additional domain-specific and location-specific pre-training on patients excluded from the study (F1-score 0·826 [95% CI 0·817–0·835], AUC 0·897 [95% CI 0·875–0·913]) performed best. All large language models, including the base model without additional pre-training (F1-score 0·739 [95% CI 0·738–0·741], AUROC 0·846 [95% CI 0·826–0·861]) outperformed models trained with structured data. With structured data only, XGBoost outperformed logistic regression and XGBoost models trained with the Boston Children’s Hospital EMR (logistic regression: F1-score 0·650 [95% CI 0·643–0·657], AUC 0·694 [95% CI 0·685–0·705], XGBoost: F1-score 0·679 [0·676–0·683], AUC 0·725 [0·717–0·734]) performed similarly to models trained on the IBM MarketScan database (logistic regression: F1-score 0·596 [0·590–0·601], AUC 0·670 [0·664–0·675], XGBoost: F1-score 0·678 [0·668–0·687], AUC 0·710 [0·703–0·714]). Physician’s clinical notes about an initial seizure-like event include substantial signals for prediction of seizure recurrence, and additional domain-specific and location-specific pre-training can significantly improve the performance of clinical large language models, even for specialised cohorts. UCB, National Institute of Neurological Disorders and Stroke (US National Institutes of Health).
DOI: 10.1016/j.seizure.2021.11.007
发表时间: 2022-01
期刊: Seizure
影响因子: --
作者:
Bonnett LJ;Kim L;Johnson A;Sander JW;Lawn N;Beghi E;Leone M;Marson AG
通讯作者: Marson AG