Grafted and Vanishing Random Subspaces.

Grafted and Vanishing Random Subspaces.
复制标题

DOI:
10.1007/s10044-021-01029-0
复制
发表时间:
2022-03
期刊:
Pattern analysis and applications : PAA
影响因子:
--
通讯作者:
Love TM
Love TM
中科院分区:
其他
文献类型:
--
作者:
Corsetti MA;Love TM

文献摘要

参考文献

相似文献

随机子空间方法 (RSM) 是一种集成过程,其中使用随机选择的数据特征子集来构造每个组成学习器。回归树是 RSM 集成中理想的候选学习器。通过根据不同的特征子集构建树,RSM 减少了树之间的相关性,从而形成更强的集成。此外,它在构建每棵树时仅考虑特征的子集,从而减轻了计算负担。尽管 RSM 有明显的优点,但它也有一个显着的缺点。在某些情况下,随机选择的子空间可能缺乏信息特征。在真正提供信息的变量数量相对于变量总数较少的情况下尤其如此。使用缺乏信息特征的特征子集构建的树可能会损害整体。在这里,我们提出了嫁接随机子空间(GRS)和消失随机子空间(VRS),这两种新颖的集成过程旨在通过跨树重用信息来弥补上述缺陷。这两种技术都借鉴了 RSM,在随机选择的特征子集上生长单独的树。对于 GRS 集成中的每棵树,都会识别出最重要的变量,并保证将其包含到接下来的 q 个特征子集中。这使得 GRS 能够在多个连续的树中循环利用一棵树中的一个有前景的特征,从而有效地将变量移植到接下来的 q 个活动子集中。在 VRS 过程中,保证将最不重要的特征从接下来的 q 个特征子集中排除。这创建了一个更丰富的候选变量池,从中提取连续的特征子集。
The Random Subspace Method (RSM) is an ensemble procedure in which each constituent learner is constructed using a randomly chosen subset of the data features. Regression trees are ideal candidate learners in RSM ensembles. By constructing trees upon different feature subsets, RSM reduces correlation between trees resulting in a stronger ensemble. Furthermore, it lessens computational burden by only considering a subset of the features when building each tree. Despite its apparent advantages, RSM has a notable drawback. In some instances a randomly chosen subspace may lack informative features. This is especially true in situations in which the number of truly informative variables is small relative to the total number of variables. Trees that are constructed using feature subsets lacking informative features can be damaging to the ensemble. Here we present Grafted Random Subspaces (GRS) and Vanishing Random Subspaces (VRS), two novel ensemble procedures designed to remedy the aforementioned drawback by reusing information across trees. Both techniques borrow from RSM by growing individual trees on randomly selected feature subsets. For each tree in a GRS ensemble, the most important variable is identified and guaranteed inclusion into the next q feature subsets. This allows GRS to recycle a promising feature from one tree across several successive trees, effectively grafting the variable into the next q active subsets. In the VRS procedure the least important feature is guaranteed exclusion from the next q feature subsets. This creates a more enriched pool of candidate variables from which the successive feature subsets are drawn.
DOI: 10.1016/0095-0696(78)90006-2
发表时间: 1978-01-01
影响因子: 4.6
作者:
HARRISON, D;RUBINFELD, DL
通讯作者: RUBINFELD, DL
DOI: 10.1198/106186006x133933
发表时间: 2006-09-01
影响因子: 2.4
作者:
Hothorn, Torsten;Hornik, Kurt;Zeileis, Achim
通讯作者: Zeileis, Achim
DOI: 10.1023/a:1022648800760
发表时间: 1990-06-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
SCHAPIRE, RE
通讯作者: SCHAPIRE, RE
DOI: 10.1016/s0031-3203(02)00121-8
发表时间: 2003-06-01
影响因子: 8
作者:
Bryll, R;Gutierrez-Osuna, R;Quek, F
通讯作者: Quek, F
DOI: 10.1061/(asce)co.1943-7862.0001047
发表时间: 2016-02-01
影响因子: 5.1
作者:
Rafiei, Mohammad Hossein;Adeli, Hojjat
通讯作者: Adeli, Hojjat