Classification in conservation biology: A comparison of five machine-learning methods

Classification in conservation biology: A comparison of five machine-learning methods
复制标题

DOI:
10.1016/j.ecoinf.2010.06.003
复制
发表时间:
2010-11-01
影响因子:
5.1
通讯作者:
Arriaga-Weiss, Stefan
Arriaga-Weiss, Stefan
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Kampichler, Christian;Wieland, Ralf;Arriaga-Weiss, Stefan

文献摘要

被引文献

相似文献

分类是生态学中应用最广泛的任务之一。生态学家必须处理嘈杂的高维数据,这些数据通常是非线性的,不符合传统统计程序的假设。为了克服这个问题,机器学习方法被用作生态分类方法。在生物保护背景下,我们比较了五种基于机器学习的分类技术(分类树、随机森林、人工神经网络、支持向量机和自动诱导基于规则的模糊模型)。该研究的案例是一种尤卡坦半岛特有的鸟类,在过去的几十年里,当地的数量和分布区域都有相当大的减少。在叠加到半岛上的10 × 10公里网格上,我们分析了1980年至2000年间环境和社会解释变量与星形火鸡丰度变化之间的关系。丰度分别表示为3个(减少、不变和增加)和14个更详细的丰度变化类。通过总体分类误差和归一化互信息指数来衡量,随机森林和分类树是最有效的方法,不同方法的建模性能差异很大。人工神经网络与线性判别分析一起产生了最差的结果,线性判别分析被包括在传统的统计方法中。我们不仅评估了分类准确性,还评估了诸如时间努力、分类器可理解性和方法复杂性等特征,这些特征决定了分类技术在生态学家和保护生物学家之间的成功,以及与管理人员和决策者的沟通。由于分类器的易解释性和方法的高可理解性,我们建议将分类树和随机森林结合使用。(C) 2010 Elsevier B.V.版权所有
Classification is one of the most widely applied tasks in ecology. Ecologists have to deal with noisy, high-dimensional data that often are non-linear and do not meet the assumptions of conventional statistical procedures. To overcome this problem, machine-learning methods have been adopted as ecological classification methods. We compared five machine-learning based classification techniques (classification trees, random forests, artificial neural networks, support vector machines, and automatically induced rule-based fuzzy models) in a biological conservation context. The study case was that of the ocellated turkey (Meleagris ocellata), a bird endemic to the Yucatan peninsula that has suffered considerable decreases in local abundance and distributional area during the last few decades. On a grid of 10x 10 km cells that was superimposed to the peninsula we analysed relationships between environmental and social explanatory variables and ocellated turkey abundance changes between 1980 and 2000. Abundance was expressed in three (decrease, no change, and increase) and 14 more detailed abundance change classes, respectively. Modelling performance varied considerably between methods with random forests and classification trees being the most efficient ones as measured by overall classification error and the normalised mutual information index. Artificial neural networks yielded the worst results along with linear discriminant analysis, which was included as a conventional statistical approach. We not only evaluated classification accuracy but also characteristics such as time effort, classifier comprehensibility and method intricacy aspects that determine the success of a classification technique among ecologists and conservation biologists as well as for the communication with managers and decision makers. We recommend the combined use of classification trees and random forests due to the easy interpretability of classifiers and the high comprehensibility of the method. (C) 2010 Elsevier B.V. All rights reserved.