Decision trees

Decision trees
复制标题

DOI:
10.1002/wics.1278
复制
发表时间:
2013-11-01
影响因子:
3.2
通讯作者:
de Ville, Barry
de Ville, Barry
中科院分区:
数学3区
文献类型:
--
作者:
de Ville, Barry

文献摘要

被引文献

相似文献

决策树的起源可以追溯到书面记录的早期发展时期。这段历史说明了树的一个主要优势:异常可解释的结果,具有直观的树状显示,这反过来又增强了对结果的理解和传播。决策树(有时称为分类树或回归树)的计算起源是生物和认知过程的模型。这种共同的传统推动了统计决策树和机器学习树的互补发展。展开和逐步阐明树木的各种特点,在整个早期的历史,在世纪后期讨论沿着重要的相关参考点和负责任的作者。统计方法,如假设检验和各种恢复方法,与机器学习实现一起沿着发展。这就产生了适应性极强的决策树工具,适用于各种统计和机器学习任务,跨越不同的计量层次,数据质量也各不相同。树在存在缺失数据的情况下是健壮的,并提供了多种将缺失数据合并到结果模型中的方法。虽然树是强大的,但它们也是灵活且易于使用的方法。这确保了高质量的结果,需要很少的假设部署的生产。最后,本文讨论了最新的发展,这些发展仍然依赖于统计和机器学习社区之间的协同作用和相互促进。目前的发展与出现的多棵树和各种rescent的方法,采用了讨论。(C)2013威利期刊公司
Decision trees trace their origins to the era of the early development of written records. This history illustrates a major strength of trees: exceptionally interpretable results which have an intuitive tree-like display which, in turn, enhances understanding and the dissemination of results. The computational origins of decision trees-sometimes called classification trees or regression trees-are models of biological and cognitive processes. This common heritage drives complementary developments of both statistical decision trees and trees designed for machine learning. The unfolding and progressive elucidation of the various features of trees throughout their early history in the late 20th century is discussed along with the important associated reference points and responsible authors. Statistical approaches, such as a hypothesis testing and various resampling approaches, have coevolved along with machine learning implementations. This had resulted in exceptionally adaptable decision tree tools, appropriate for various statistical and machine learning tasks, across various levels of measurement, with varying levels of data quality. Trees are robust in the presence of missing data and offer multiple ways of incorporating missing data in the resulting models. Although trees are powerful, they are also flexible and easy to use methods. This assures the production of high quality results that require few assumptions to deploy. The treatment ends with a discussion of the most current developments which continue to rely on the synergies and cross-fertilization between statistical and machine learning communities. Current developments with the emergence of multiple trees and the various resampling approaches that are employed are discussed. (C) 2013 Wiley Periodicals, Inc.