Mapping ecological systems with a random forest model: tradeoffs between errors and bias

Mapping ecological systems with a random forest model: tradeoffs between errors and bias
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
Emilie B. Grossmann;J. Ohmann;James S. Kagan;H. May;M. Gregory
Emilie B. Grossmann;J. Ohmann;James S. Kagan;H. May;M. Gregory
中科院分区:
其他
文献类型:
--
作者:
Emilie B. Grossmann;J. Ohmann;James S. Kagan;H. May;M. Gregory

文献摘要

被引文献

相似文献

自 LANDSAT 计划发起以来的过去三十年里,利用遥感数据生成植被地图的方法取得了巨大进步。它们的范围从单幅图像的监督和无监督分类,到基于多时相图像的分类,再到辅助信息与遥感图像的集成(Holmgren 和Thuresson 1998)。后一种技术可以更详细、更准确地估计植物群落组成,它们对于创建 USGS GAP 植被层 2000 年更新至关重要。 Nature Serve 的生态系统(系统)在物种组成方面的特异性水平使得其中许多无法仅根据图像进行区分。这是植被遥感的一个常见问题(Kalliola 和 Syrjanen 1991)。然而,气候、地貌和土壤的图像和辅助信息的结合通常可以提供足够的信息,以 30m 的分辨率绘制整个景观的系统地图。分类树 (CART) 及其扩展是一系列常用于生态分析的建模技术(De'ath 和 Fabricius 2000,Cutler 等人 2007)。 CART 模型还用于根据植被、图像和辅助环境数据之间的关系构建预测植被地图(例如,Franklin 2002)。单 CART 模型是通过递归划分构建的,其中响应变量按顺序迭代地划分为组,组“纯度”随着每次划分而增加(Breiman 等人,1984)。划分是基于解释变量内的阈值。 CART 模型已通过 ERDAS Imagine 软件的 See5/C5.0 模块在绘图中得到普及,并已用于构建其他地区的 GAP 植被层(Lowry 2005)。然而,CART 模型容易过度拟合数据,从而导致预测错误。随机森林 (RF) 模型是 CART 的扩展,可限制过度拟合问题。 RF 不是根据所有可用数据构建单个预测树模型,而是构建数百个树模型,使用随机的绘图数据子集和解释变量来构建每棵树。这种内部交叉验证过程可以防止单个 CART 模型固有的过度拟合问题 (Breiman 2001),因此它们在植被绘图中变得越来越流行 (Prasad et al. 2006, Iverson et al. 2008, Evans and Cushman 2009)。然而,RF 模型可能会出现偏差问题,尤其是当类之间的绘图样本不平衡时(Chen 等人,2004 年)。由于生态系统很少在任何给定区域占据相同的面积,因此它们在系统样地样本中的表示通常是不平衡的。在我们绘制多分辨率土地特征联盟 (MRLC) 地图区域 2 和 7 的工作中,我们探索了 GAP 绘图过程中绘图方法的影响,重点关注 RF 作为一种有前景的技术,因为它以从噪声、非正态数据中进行准确的分类预测而闻名 (Breiman 2001)。在这里,我们展示了俄勒冈州西部西喀斯喀特生态区森林生态系统的两张对比图(图 1),基于:a) RF 和 b) RF 以及相关的偏差调整程序 (RF_Adj)。我们对比了它们的差异、优点和缺点,并为未来的 GAP 植被制图工作提出了一些建议。请注意,此处提供的地图并非最终的 GAP
Methods for generating vegetation maps from remotely sensed data have advanced greatly within the last three decades since the LANDSAT program originated. They range from supervised and unsupervised classifications of single images, to classifications based on multitemporal imagery, to the integration of ancillary information with remotely sensed imagery (Holmgren and Thuresson 1998). The latter techniques allow more detailed and accurate estimations of plant community composition and they were essential for creating the 2000 update for the USGS GAP vegetation layer. The level of specificity of Nature Serve's Ecological Systems (Systems) with respect to species composition makes many of them impossible to differentiate based on imagery alone. This is a common problem with remote sensing of vegetation (Kalliola and Syrjanen 1991). However, the combination of imagery and ancillary information on climate, landform and soil often provides enough information to map the Systems across the landscape at 30m resolution. Classification trees (CART) and their extensions are a family of modeling techniques that are often used in ecological analysis (De'ath and Fabricius 2000, Cutler et al. 2007). CART models are also used to build predictive vegetation maps, based on relationships between vegetation, imagery and ancillary environmental data (e.g., Franklin 2002). Single CART models are built through recursive partitioning, wherein the response variable is iteratively divided into groups sequentially with group 'purity' increasing with each division (Breiman et al. 1984). Divisions are based on thresholds within explanatory variables. CART models have been popularized for mapping through the See5/C5.0 module for ERDAS Imagine software, and have been used to build the GAP vegetation layer in other regions (Lowry 2005). CART models, however, are prone to overfitting data, which can lead to predictive errors. Random forest (RF) models are an extension of CART that limits the over-fitting problem. Rather than building a single predictive tree model from all available data, RF builds hundreds of tree models, using randomized subsets of plot data and explanatory variables to build each tree. This process of internal cross-validation prevents the over-fitting problem inherent to a single CART model (Breiman 2001), hence they are becoming more popular for vegetation mapping (Prasad et al. 2006, Iverson et al. 2008, Evans and Cushman 2009). However, RF models can exhibit bias problems especially when plot-samples are unbalanced among the classes (Chen et al. 2004). Because Ecological Systems seldom occupy equal areas across any given region, their representation within systematic plot samples is normally unbalanced. In our work mapping Multi-Resolution Land Characteristics Consortium (MRLC) mapzones 2 and 7, we explored the implications of mapping methods in the GAP mapping process, focusing on RF as a promising technique because it is known for making accurate classification predictions from noisy, non-normal data (Breiman 2001). Here, we present two contrasting maps of forested Ecological Systems across the West Cascades ecoregion (Figure 1) in Western Oregon based on: a) RF and b) RF with an associated bias adjustment procedure (RF_Adj). We contrast their differences, strengths and weaknesses, and make some recommendations for future GAP vegetation mapping efforts. Note that the maps presented here are not final GAP