Mapping ecological systems with a random forest model: tradeoffs between errors and bias
Mapping ecological systems with a random forest model: tradeoffs between errors and bias
复制标题
DOI:
--
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
Emilie B. Grossmann;J. Ohmann;James S. Kagan;H. May;M. Gregory
中科院分区:
文献类型:
--
作者:
Emilie B. Grossmann;J. Ohmann;James S. Kagan;H. May;M. Gregory
Methods for generating vegetation maps from remotely sensed data have advanced greatly within the last three decades since the LANDSAT program originated. They range from supervised and unsupervised classifications of single images, to classifications based on multitemporal imagery, to the integration of ancillary information with remotely sensed imagery (Holmgren and Thuresson 1998). The latter techniques allow more detailed and accurate estimations of plant community composition and they were essential for creating the 2000 update for the USGS GAP vegetation layer. The level of specificity of Nature Serve's Ecological Systems (Systems) with respect to species composition makes many of them impossible to differentiate based on imagery alone. This is a common problem with remote sensing of vegetation (Kalliola and Syrjanen 1991). However, the combination of imagery and ancillary information on climate, landform and soil often provides enough information to map the Systems across the landscape at 30m resolution. Classification trees (CART) and their extensions are a family of modeling techniques that are often used in ecological analysis (De'ath and Fabricius 2000, Cutler et al. 2007). CART models are also used to build predictive vegetation maps, based on relationships between vegetation, imagery and ancillary environmental data (e.g., Franklin 2002). Single CART models are built through recursive partitioning, wherein the response variable is iteratively divided into groups sequentially with group 'purity' increasing with each division (Breiman et al. 1984). Divisions are based on thresholds within explanatory variables. CART models have been popularized for mapping through the See5/C5.0 module for ERDAS Imagine software, and have been used to build the GAP vegetation layer in other regions (Lowry 2005). CART models, however, are prone to overfitting data, which can lead to predictive errors. Random forest (RF) models are an extension of CART that limits the over-fitting problem. Rather than building a single predictive tree model from all available data, RF builds hundreds of tree models, using randomized subsets of plot data and explanatory variables to build each tree. This process of internal cross-validation prevents the over-fitting problem inherent to a single CART model (Breiman 2001), hence they are becoming more popular for vegetation mapping (Prasad et al. 2006, Iverson et al. 2008, Evans and Cushman 2009). However, RF models can exhibit bias problems especially when plot-samples are unbalanced among the classes (Chen et al. 2004). Because Ecological Systems seldom occupy equal areas across any given region, their representation within systematic plot samples is normally unbalanced. In our work mapping Multi-Resolution Land Characteristics Consortium (MRLC) mapzones 2 and 7, we explored the implications of mapping methods in the GAP mapping process, focusing on RF as a promising technique because it is known for making accurate classification predictions from noisy, non-normal data (Breiman 2001). Here, we present two contrasting maps of forested Ecological Systems across the West Cascades ecoregion (Figure 1) in Western Oregon based on: a) RF and b) RF with an associated bias adjustment procedure (RF_Adj). We contrast their differences, strengths and weaknesses, and make some recommendations for future GAP vegetation mapping efforts. Note that the maps presented here are not final GAP