Machine learning approaches to the social determinants of health in the health and retirement study.

Machine learning approaches to the social determinants of health in the health and retirement study.
复制标题

DOI:
10.1016/j.ssmph.2017.11.008
复制
发表时间:
2018-04
期刊:
SSM - population health
影响因子:
--
通讯作者:
Rehkopf D
Rehkopf D
中科院分区:
其他
文献类型:
--
作者:
Seligman B;Tuljapurkar S;Rehkopf D

文献摘要

相似文献

社会和经济因素是健康的重要预测因素,对卫生系统具有公认的重要性。然而,生物医学文献中其他地方使用的机器学习并没有被广泛应用于研究社会和健康之间的关系。我们使用来自健康与退休研究的数据,研究机器学习如何增加我们对健康的社会决定因素的理解。年龄和性别的线性回归,以及加上收入、财富和教育的基于简约理论的回归,被用来预测收缩压、体重指数、腰围和端粒长度。我们比较了四种机器学习方法:线性回归、惩罚回归、随机森林和神经网络的预测性、匹配性和可解释性。所有模型的样本外预测都很差。大多数机器学习模型的表现与更简单的模型相似。然而,神经网络的表现远远好于其他三种方法。神经网络对数据的拟合也很好(R2在0.4-0.6之间,而其他所有数据都是0.3)。在机器学习模型中,九个变量经常被选为预测变量或权重较高的变量:牙科就诊、当前吸烟、自我评估健康状况、连续七次减法、继承遗产的可能性、留下至少1万美元遗产的可能性、出生的孩子数量、非裔美国人种族和性别。一些机器学习方法不能改善预测,也不适合更简单的模型,然而,神经网络表现得很好。跨模型确定的预测因素表明,潜在的社会因素是慢性病生物学指标的重要预测因素,而神经网络方法的基础变量之间的非线性和交互关系可能是重要的考虑因素。“大数据”方法可能有助于理解健康的社会决定因素。神经网络在预测和解释方差方面优于其他方法。没有一种单独的机器学习方法是容易解释的。机器学习方法中共有的变量表明了核心的社会决定因素。
Social and economic factors are important predictors of health and of recognized importance for health systems. However, machine learning, used elsewhere in the biomedical literature, has not been extensively applied to study relationships between society and health. We investigate how machine learning may add to our understanding of social determinants of health using data from the Health and Retirement Study. A linear regression of age and gender, and a parsimonious theory-based regression additionally incorporating income, wealth, and education, were used to predict systolic blood pressure, body mass index, waist circumference, and telomere length. Prediction, fit, and interpretability were compared across four machine learning methods: linear regression, penalized regressions, random forests, and neural networks. All models had poor out-of-sample prediction. Most machine learning models performed similarly to the simpler models. However, neural networks greatly outperformed the three other methods. Neural networks also had good fit to the data (R2 between 0.4–0.6, versus <0.3 for all others). Across machine learning models, nine variables were frequently selected or highly weighted as predictors: dental visits, current smoking, self-rated health, serial-seven subtractions, probability of receiving an inheritance, probability of leaving an inheritance of at least $10,000, number of children ever born, African-American race, and gender. Some of the machine learning methods do not improve prediction or fit beyond simpler models, however, neural networks performed well. The predictors identified across models suggest underlying social factors that are important predictors of biological indicators of chronic disease, and that the non-linear and interactive relationships between variables fundamental to the neural network approach may be important to consider. “Big data” methods may aid understanding of social determinants of health. Neural networks outperform other methods in prediction and variance explained. No individual machine learning method was readily interpretable. Variables in common among machine learning methods suggest core social determinants.