Unsupervised Learning
Unsupervised Learning
复制标题
DOI:
10.1007/978-1-4614-7138-7_10
复制
发表时间:
2013-01-01
期刊:
影响因子:
--
通讯作者:
Tibshirani, Robert
中科院分区:
文献类型:
--
作者:
James, Gareth;Witten, Daniela;Tibshirani, Robert
Most of this book concerns supervised learning methods such as regression and classification. In the supervised learning setting, we typically have access to a set of p features X1, X2,..., Xp, measured on n observations, and a response Y also measured on those same n observations. The goal is then to predict Y using X1, X2,..., Xp. This chapter will instead focus on unsupervised learning, a set of statistical tools intended for the setting in which we have only a set of features X1, X2,..., Xp measured on n observations. We are not interested in prediction, because we do not have an associated response variable Y. Rather, the goal is to discover interesting things about the measurements on X1, X2,..., Xp. Is there an informative way to visualize the data? Can we discover subgroups among the variables or among the observations? Unsupervised learning refers to a diverse set of techniques for answering questions such as these. In this chapter, we will focus on two particular types of unsupervised learning: principal components analysis, a tool used for data visualization or data pre-processing before supervised techniques are applied, and clustering, a broad class of methods for discovering unknown subgroups in data.