Prediction model development of late-onset preeclampsia using machine learning-based methods

Prediction model development of late-onset preeclampsia using machine learning-based methods
复制标题

DOI:
10.1371/journal.pone.0221202
复制
发表时间:
2019-08-23
期刊:
影响因子:
3.7
通讯作者:
Park, Jung Tak
Park, Jung Tak
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Jhee, Jong Hyun;Lee, SungHee;Park, Jung Tak

文献摘要

被引文献

相似文献

先兆子痫是孕产妇和胎儿发病和死亡的主要原因之一。由于缺乏有效的预防措施,其预测对于及时管理至关重要。本研究旨在开发利用机器学习的模型,利用医院电子病历数据来预测迟发性先兆子痫。还比较了基于机器学习的模型和使用传统统计方法的模型的性能。共有 11,006 名在延世大学医院接受产前护理的孕妇纳入其中。从妊娠中期早期至 34 周期间的电子病历中检索孕产妇数据。预测结果是妊娠 34 周后发生晚发性先兆子痫。使用模式识别和聚类分析来选择预测模型中包含的参数。使用逻辑回归、决策树模型、朴素贝叶斯分类、支持向量机、随机森林算法和随机梯度提升方法来构建预测模型。 C 统计量用于评估每个模型的性能。先兆子痫的总体发生率为 4.7%(474 名患者)。收缩压、血清尿素氮和肌酐水平、血小板计数、血清钾水平、白细胞计数、血清钙水平和尿蛋白是预测模型中最有影响力的变量。决策树模型、朴素贝叶斯分类、支持向量机、随机森林算法、随机梯度提升方法和逻辑回归模型的 C 统计量分别为 0.857、0.776、0.573、0.894、0.924 和 0.806。随机梯度增强模型的预测性能最好,准确率和误报率分别为 0.973 和 0.009。结合使用母体因素和妊娠中期早期至妊娠晚期的常见产前实验室数据,可以使用机器学习算法有效预测迟发性先兆子痫。需要未来的前瞻性研究来验证算法的临床适用性。
Preeclampsia is one of the leading causes of maternal and fetal morbidity and mortality. Due to the lack of effective preventive measures, its prediction is essential to its prompt management. This study aimed to develop models using machine learning to predict late-onset preeclampsia using hospital electronic medical record data. The performance of the machine learning based models and models using conventional statistical methods were also compared. A total of 11,006 pregnant women who received antenatal care at Yonsei University Hospital were included. Maternal data were retrieved from electronic medical records during the early second trimester to 34 weeks. The prediction outcome was late-onset preeclampsia occurrence after 34 weeks' gestation. Pattern recognition and cluster analysis were used to select the parameters included in the prediction models. Logistic regression, decision tree model, naive Bayes classification, support vector machine, random forest algorithm, and stochastic gradient boosting method were used to construct the prediction models. C-statistics was used to assess the performance of each model. The overall preeclampsia development rate was 4.7% (474 patients). Systolic blood pressure, serum blood urea nitrogen and creatinine levels, platelet counts, serum potassium level, white blood cell count, serum calcium level, and urinary protein were the most influential variables included in the prediction models. C-statistics for the decision tree model, naive Bayes classification, support vector machine, random forest algorithm, stochastic gradient boosting method, and logistic regression models were 0.857, 0.776, 0.573, 0.894, 0.924, and 0.806, respectively. The stochastic gradient boosting model had the best prediction performance with an accuracy and false positive rate of 0.973 and 0.009, respectively. The combined use of maternal factors and common antenatal laboratory data of the early second trimester through early third trimester could effectively predict late-onset preeclampsia using machine learning algorithms. Future prospective studies are needed to verify the clinical applicability algorithms.