A Machine Learning Approach to College Drinking Prediction and Risk Factor Identification

A Machine Learning Approach to College Drinking Prediction and Risk Factor Identification
复制标题

DOI:
10.1145/2508037.2508053
复制
发表时间:
2013-09-01
影响因子:
5
通讯作者:
Armeli, Stephen
Armeli, Stephen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Bi, Jinbo;Sun, Jiangwen;Armeli, Stephen

文献摘要

被引文献

相似文献

酒精滥用是美国青少年和年轻人面临的最严重的公共卫生问题之一。国家统计数据显示,21岁以下青年饮酒中有近90%涉及酗酒,44%的大学生从事高风险饮酒活动。传统的酒精干预计划,旨在建立一个酒精减少规范或禁止未成年人饮酒,多年来在控制大学酗酒方面几乎没有取得进展。现有的酒精研究是演绎性的,收集数据以调查心理/行为假设,并对数据进行统计分析以确认假设。由于这种验证性的分析方式,所得到的统计模型是群体特异性的,并且通常无法在不同的样品上重复。本文介绍了两种机器学习方法,用于对美国国家酒精滥用和酒精中毒研究所赞助的大学酒精研究中收集的纵向数据进行二次分析。我们的方法旨在从多波队列连续的每日数据中发现知识,这些数据可能与原始假设一致,也可能不一致,但量化预测模型的可能性更高,可以推广到新的样本。首先,我们提出了一个所谓的时间相关的支持向量机构建一个分类器作为日常情绪,压力和饮酒预期的函数,以区分夜间狂饮的日子,没有个别学生。然后,我们提出了一个组合的聚类分析和特征选择,聚类分析是用来识别饮酒模式的基础上,平均每日饮酒行为和特征选择是用来识别与每个模式相关的风险因素。我们评估我们的方法,在春季和秋季学期,分别招募了两个队列的530名大学生。这两个队列和进一步的100个随机分区的总学生的交叉验证表明,我们的方法提高了模型的泛化能力相比,传统的多层次逻辑回归。在我们的模型中发现的风险因素和这些因素之间的相互作用可以为更有效的大学酒精干预的新设计奠定潜在的基础并提供见解。
Alcohol misuse is one of the most serious public health problems facing adolescents and young adults in the United States. National statistics shows that nearly 90% of alcohol consumed by youth under 21 years of age involves binge drinking and 44% of college students engage in high-risk drinking activities. Conventional alcohol intervention programs, which aim at installing either an alcohol reduction norm or prohibition against underage drinking, have yielded little progress in controlling college binge drinking over the years. Existing alcohol studies are deductive where data are collected to investigate a psychological/behavioral hypothesis, and statistical analysis is applied to the data to confirm the hypothesis. Due to this confirmatory manner of analysis, the resulting statistical models are cohort-specific and typically fail to replicate on a different sample. This article presents two machine learning approaches for a secondary analysis of longitudinal data collected in college alcohol studies sponsored by the National Institute on Alcohol Abuse and Alcoholism. Our approach aims to discover knowledge, from multiwave cohort-sequential daily data, which may or may not align with the original hypothesis but quantifies predictive models with higher likelihood to generalize to new samples. We first propose a so-called temporally-correlated support vector machine to construct a classifier as a function of daily moods, stress, and drinking expectancies to distinguish days with nighttime binge drinking from days without for individual students. We then propose a combination of cluster analysis and feature selection, where cluster analysis is used to identify drinking patterns based on averaged daily drinking behavior and feature selection is used to identify risk factors associated with each pattern. We evaluate our methods on two cohorts of 530 total college students recruited during the Spring and Fall semesters, respectively. Cross validation on these two cohorts and further on 100 random partitions of the total students demonstrate that our methods improve the model generalizability in comparison with traditional multilevel logistic regression. The discovered risk factors and the interaction of these factors delineated in our models can set a potential basis and offer insights to a new design of more effective college alcohol interventions.