Unsupervised learning with random forest predictors

Unsupervised learning with random forest predictors
复制标题

DOI:
10.1198/106186006x94072
复制
发表时间:
2006-03-01
影响因子:
2.4
通讯作者:
Horvath, S
Horvath, S
中科院分区:
数学2区
文献类型:
--
作者:
Shi, T;Horvath, S

文献摘要

被引文献

相似文献

随机森林 (RF) 预测器是单个树预测器的集合。作为其构建的一部分,RF 预测器自然会导致观测值之间的差异测量。人们还可以定义未标记数据之间的 RF 相异性度量:其想法是构建一个 RF 预测器,将“观察到的”数据与适当生成的合成数据区分开来。观察到的数据是原始的未标记数据,合成数据是从参考分布中提取的。在这里,我们描述了 RF 相异性的属性,并就如何在实践中使用它提出了建议。 RF 相异性可能很有吸引力,因为它可以很好地处理混合变量类型,对于输入变量的单调变换具有不变性,并且对于外围观测值具有鲁棒性。 RF相异性由于其内在的变量选择,可以轻松处理大量变量;例如,Addcl1 RF 相异性根据每个变量对其他变量的依赖程度来衡量每个变量的贡献。我们发现 RF 相异性对于基于肿瘤标志物表达来检测肿瘤样本簇很有用。在此应用中,具有生物学意义的簇通常可以用简单的阈值规则来描述。
A random forest (RF) predictor is an ensemble of individual tree predictors. As part of their construction, RF predictors naturally lead to a dissimilarity measure between the observations. One can also define an RF dissimilarity measure between unlabeled data: the idea is to construct an RF predictor that distinguishes the "observed" data from suitably generated synthetic data. The observed data are the original unlabeled data and the synthetic data are drawn from a reference distribution. Here we describe the properties of the RF dissimilarity and make recommendations on how to use it in practice.An RF dissimilarity can be attractive because it handles mixed variable types well, is invariant to monotonic transformations of the input variables, and is robust to outlying observations. The RF dissimilarity easily deals with a large number of variables due to its intrinsic variable selection; for example, the Addcl1 RF dissimilarity weighs the contribution of each variable according to how dependent it is on other variables.We find that the RF dissimilarity is useful for detecting tumor sample clusters on the basis of tumor marker expressions. In this application, biologically meaningful clusters can often be described with simple thresholding rules.