A comparative study: classification vs. user-based collaborative filtering for clinical prediction.
A comparative study: classification vs. user-based collaborative filtering for clinical prediction.
复制标题
DOI:
10.1186/s12874-016-0261-9
复制
发表时间:
2016-12-08
影响因子:
4
通讯作者:
Blair RH
中科院分区:
文献类型:
--
作者:
Hao F;Blair RH
Recommender systems have shown tremendous value for the prediction of personalized item recommendations for individuals in a variety of settings (e.g., marketing, e-commerce, etc.). User-based collaborative filtering is a popular recommender system, which leverages an individuals’ prior satisfaction with items, as well as the satisfaction of individuals that are “similar”. Recently, there have been applications of collaborative filtering based recommender systems for clinical risk prediction. In these applications, individuals represent patients, and items represent clinical data, which includes an outcome. Application of recommender systems to a problem of this type requires the recasting a supervised learning problem as unsupervised. The rationale is that patients with similar clinical features carry a similar disease risk. As the “Big Data” era progresses, it is likely that approaches of this type will be reached for as biomedical data continues to grow in both size and complexity (e.g., electronic health records). In the present study, we set out to understand and assess the performance of recommender systems in a controlled yet realistic setting. User-based collaborative filtering recommender systems are compared to logistic regression and random forests with different types of imputation and varying amounts of missingness on four different publicly available medical data sets: National Health and Nutrition Examination Survey (NHANES, 2011-2012 on Obesity), Study to Understand Prognoses Preferences Outcomes and Risks of Treatment (SUPPORT), chronic kidney disease, and dermatology data. We also examined performance using simulated data with observations that are Missing At Random (MAR) or Missing Completely At Random (MCAR) under various degrees of missingness and levels of class imbalance in the response variable. Our results demonstrate that user-based collaborative filtering is consistently inferior to logistic regression and random forests with different imputations on real and simulated data. The results warrant caution for the collaborative filtering for the purpose of clinical risk prediction when traditional classification is feasible and practical. CF may not be desirable in datasets where classification is an acceptable alternative. We describe some natural applications related to “Big Data” where CF would be preferred and conclude with some insights as to why caution may be warranted in this context. The online version of this article (doi:10.1186/s12874-016-0261-9) contains supplementary material, which is available to authorized users.
登录
查看更多内容
影响因子:
1.8
作者:
Heitjan, DF;Basu, S
通讯作者:
Basu, S
DOI:
10.1136/amiajnl-2014-002974
发表时间:
2014-11
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
作者:
Margolis R;Derr L;Dunn M;Huerta M;Larkin J;Sheehan J;Guyer M;Green ED
通讯作者:
Green ED
影响因子:
4.8
作者:
Davis, Darcy A.;Chawla, Nitesh V.;Barabasi, Albert-Laszlo
通讯作者:
Barabasi, Albert-Laszlo
影响因子:
37.8
作者:
Scirica, Benjamin M.;Morrow, David A.;Braunwald, Eugene
通讯作者:
Braunwald, Eugene
DOI:
10.1080/03610928008827941
发表时间:
1980-01-01
期刊:
COMMUNICATIONS IN STATISTICS PART A-THEORY AND METHODS
影响因子:
--
作者:
HOSMER, DW;LEMESHOW, S
通讯作者:
LEMESHOW, S