A Prediction Model for Uncontrolled Type 2 Diabetes Mellitus Incorporating Area-level Social Determinants of Health

A Prediction Model for Uncontrolled Type 2 Diabetes Mellitus Incorporating Area-level Social Determinants of Health
复制标题

DOI:
10.1097/mlr.0000000000001147
复制
发表时间:
2019-08-01
期刊:
影响因子:
3
通讯作者:
Narayanaswamy, Rajiv
Narayanaswamy, Rajiv
中科院分区:
医学3区
文献类型:
--
作者:
Basu, Sanjay;Narayanaswamy, Rajiv

文献摘要

被引文献

相似文献

背景资料:在地区层面,健康的社会决定因素(SDH)被认为会影响2型糖尿病(T2 DM)患者血糖控制不良的可能性。目的:建立一个模型,预测T2 DM患者是否患有未控制的糖尿病(血红蛋白A1 c>= 9%),纳入个人和地区水平(人口普查区)协变量。研究设计:机器学习模型的开发和验证。受试者:索赔数据中共有N= 1,015,808名T2 DM私人投保人。测量:C统计量、灵敏度、特异性、阳性预测值、阴性预测值和准确性。结果如下:在可用的个人水平协变量和地区水平SDH协变量中选择的标准logistic回归模型(在人口普查区水平)表现不佳,C-统计量为0.685,灵敏度为25.6%,特异性为90.1%,阳性预测值为56.9%,阴性预测值为70.4%,和68.4%的准确性,对25%的数据保持出验证子集。相比之下,机器学习模型在风险预测方面有所改善,随机森林算法的性能最高,C统计量为0.928,灵敏度为68.5%,特异性为94.6%,阳性预测值为69.8%,阴性预测值为94.3%,准确率为90.6%。SDH变量单独解释了16.9%的未控制糖尿病的变异。结论:通过机器学习方法开发的预测模型可以帮助医疗保健组织确定要监测哪些地区级SDH数据以预测糖尿病控制,从而可能用于风险调整和目标确定。
Background: Social determinants of health (SDH) at the area level are understood to influence the likelihood of having poor glycemic control for patients with type 2 diabetes mellitus (T2DM). Objectives: To develop a model for predicting whether a person with T2DM has uncontrolled diabetes (hemoglobin A1c >= 9%), incorporating individual and area-level (census tract) covariates. Research Design: Development and validation of machine learning models. Subjects: Total of N=1,015,808 privately insured persons in claims data with T2DM. Measures: C-statistic, sensitivity, specificity, positive predictive value, negative predictive value, and accuracy. Results: A standard logistic regression model selecting among the available individual-level covariates and area-level SDH covariates (at the census tract level) performed poorly, with a C-statistic of 0.685, sensitivity of 25.6%, specificity of 90.1%, positive predictive value of 56.9%, negative predictive value of 70.4%, and accuracy of 68.4% on a 25% held-out validation subset of the data. By contrast, machine learning models improved upon risk prediction, with the highest performance from a random forest algorithm with a C-statistic of 0.928, sensitivity of 68.5%, specificity of 94.6%, positive predictive value of 69.8%, negative predictive value of 94.3%, and accuracy of 90.6%. SDH variables alone explained 16.9% of variation in uncontrolled diabetes. Conclusions: A predictive model developed through a machine learning approach may assist health care organizations to identify which area-level SDH data to monitor for prediction of diabetes control, for potential use in risk-adjustment and targeting.