Creating an automated trigger for sepsis clinical decision support at emergency department triage using machine learning.

Creating an automated trigger for sepsis clinical decision support at emergency department triage using machine learning.
复制标题

DOI:
10.1371/journal.pone.0174708
复制
发表时间:
2017
期刊:
影响因子:
3.7
通讯作者:
Nathanson LA
Nathanson LA
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Horng S;Sontag DA;Halpern Y;Jernite Y;Shapiro NI;Nathanson LA

文献摘要

被引文献

相似文献

为了证明使用免费文本数据以及生命体征和人口统计数据来识别急诊科疑似感染患者的增量效益。这是一项在某三级教学医院进行的回顾性、观察性队列研究。所有在08年12月17日至13年2月17日期间连续就诊的急诊科患者均被纳入研究。没有患者被排除在外。主要结局指标是在急诊科诊断的感染,定义为患者有感染相关的ED ICD-9-CM出院诊断。患者被随机分配到训练(64%)、验证(20%)和测试(16%)数据集。在使用双元图和阴性检测对自由文本进行预处理后,我们建立了四个模型来预测感染,增量添加生命体征,主诉和自由文本护理评估。我们使用了两种不同的方法来表示自由文本:一个词袋模型和一个主题模型。然后使用支持向量机建立预测模型。我们计算了受者工作特征曲线下的面积来比较每个模型的区分能力。该研究共纳入了230,936例患者就诊。大约14%的患者的主要结局是确诊感染。仅使用生命体征和人口统计学数据的生命体征模型的ROC曲线下面积(AUC)在训练数据集为0.67,验证数据集为0.67,测试数据集为0.67 (95% CI 0.65-0.69)。包括人口统计和生命体征数据的主诉模型的AUC在训练数据集为0.84,验证数据集为0.83,测试数据集为0.83 (95% CI 0.81-0.84)。最好的方法是利用所有的自由文本。特别是,训练数据集的词袋模型的AUC为0.89,验证数据集的AUC为0.86,测试数据集的AUC为0.86 (95% CI 0.85-0.87)。训练数据集的主题模型的AUC为0.86,验证数据集的AUC为0.86,测试数据集的AUC为0.85 (95% CI 0.84-0.86)。与之前仅使用结构化数据(如生命体征和人口统计信息)的工作相比,使用自由文本大大提高了识别感染的区分能力(AUC从0.67增加到0.86)。
To demonstrate the incremental benefit of using free text data in addition to vital sign and demographic data to identify patients with suspected infection in the emergency department. This was a retrospective, observational cohort study performed at a tertiary academic teaching hospital. All consecutive ED patient visits between 12/17/08 and 2/17/13 were included. No patients were excluded. The primary outcome measure was infection diagnosed in the emergency department defined as a patient having an infection related ED ICD-9-CM discharge diagnosis. Patients were randomly allocated to train (64%), validate (20%), and test (16%) data sets. After preprocessing the free text using bigram and negation detection, we built four models to predict infection, incrementally adding vital signs, chief complaint, and free text nursing assessment. We used two different methods to represent free text: a bag of words model and a topic model. We then used a support vector machine to build the prediction model. We calculated the area under the receiver operating characteristic curve to compare the discriminatory power of each model. A total of 230,936 patient visits were included in the study. Approximately 14% of patients had the primary outcome of diagnosed infection. The area under the ROC curve (AUC) for the vitals model, which used only vital signs and demographic data, was 0.67 for the training data set, 0.67 for the validation data set, and 0.67 (95% CI 0.65–0.69) for the test data set. The AUC for the chief complaint model which also included demographic and vital sign data was 0.84 for the training data set, 0.83 for the validation data set, and 0.83 (95% CI 0.81–0.84) for the test data set. The best performing methods made use of all of the free text. In particular, the AUC for the bag-of-words model was 0.89 for training data set, 0.86 for the validation data set, and 0.86 (95% CI 0.85–0.87) for the test data set. The AUC for the topic model was 0.86 for the training data set, 0.86 for the validation data set, and 0.85 (95% CI 0.84–0.86) for the test data set. Compared to previous work that only used structured data such as vital signs and demographic information, utilizing free text drastically improves the discriminatory ability (increase in AUC from 0.67 to 0.86) of identifying infection.