Unsupervised Learning

Unsupervised Learning
复制标题

DOI:
10.1007/978-1-4614-7138-7_10
复制
发表时间:
2013-01-01
期刊:
INTRODUCTION TO STATISTICAL LEARNING: WITH APPLICATIONS IN R
影响因子:
--
通讯作者:
Tibshirani, Robert
Tibshirani, Robert
中科院分区:
其他
文献类型:
--
作者:
James, Gareth;Witten, Daniela;Tibshirani, Robert

文献摘要

被引文献

相似文献

本书的大部分内容都涉及监督学习方法,如回归和分类。在监督学习设置中,我们通常可以访问p个特征X1,X2,.,Xp,在n个观测值上测量,响应Y也在相同的n个观测值上测量。然后,目标是使用X1,X2,...,XP.本章将重点讨论无监督学习,这是一组统计工具,适用于我们只有一组特征X1,X2,.,在n个观测值上测量的Xp。我们对预测不感兴趣,因为我们没有相关的响应变量Y。相反,目标是发现关于X1,X2,.,XP.是否有一种信息化的方式来可视化数据?我们能在变量或观测值中发现子群吗?无监督学习指的是一组不同的技术来回答这些问题。在本章中,我们将重点介绍两种特殊类型的无监督学习:主成分分析,一种在应用监督技术之前用于数据可视化或数据预处理的工具,以及聚类,一种用于发现数据中未知子组的广泛方法。
Most of this book concerns supervised learning methods such as regression and classification. In the supervised learning setting, we typically have access to a set of p features X1, X2,..., Xp, measured on n observations, and a response Y also measured on those same n observations. The goal is then to predict Y using X1, X2,..., Xp. This chapter will instead focus on unsupervised learning, a set of statistical tools intended for the setting in which we have only a set of features X1, X2,..., Xp measured on n observations. We are not interested in prediction, because we do not have an associated response variable Y. Rather, the goal is to discover interesting things about the measurements on X1, X2,..., Xp. Is there an informative way to visualize the data? Can we discover subgroups among the variables or among the observations? Unsupervised learning refers to a diverse set of techniques for answering questions such as these. In this chapter, we will focus on two particular types of unsupervised learning: principal components analysis, a tool used for data visualization or data pre-processing before supervised techniques are applied, and clustering, a broad class of methods for discovering unknown subgroups in data.