Statistical modeling: The two cultures

Statistical modeling: The two cultures
复制标题

DOI:
10.1214/ss/1009213726
复制
发表时间:
2001-08-01
影响因子:
5.7
通讯作者:
Breiman, L
Breiman, L
中科院分区:
数学2区
文献类型:
--
作者:
Breiman, L

文献摘要

被引文献

相似文献

在使用统计建模从数据中得出结论方面有两种文化。假设数据是由给定的随机数据模型生成的。另一种使用算法模型,并将数据机制视为未知。统计界一直致力于几乎只使用数据模型。这种承诺导致了不相关的理论,可疑的结论,并使统计学家无法研究大量有趣的当前问题。算法建模在理论和实践上都在统计学以外的领域得到了迅速的发展。它既可以用于大型复杂数据集,也可以作为对较小数据集进行数据建模的更准确、信息更丰富的替代方法。如果我们作为一个领域的目标是使用数据来解决问题,那么我们需要摆脱对数据模型的独家依赖,并采用更多样化的工具集。
There are two cultures in the use of statistical modeling to reach conclusions from data. One assumes that the data are generated by a given stochastic data model. The other uses algorithmic models and treats the data mechanism as unknown. The statistical community has been committed to the almost exclusive use of data models. This commitment has led to irrelevant theory, questionable conclusions, and has kept statisticians from working on a large range of interesting current problems. Algorithmic modeling, both in theory and practice, has developed rapidly in fields outside statistics. It can be used both on large complex data sets and as a more accurate and informative alternative to data modeling on smaller data sets. If our goal as a field is to use data to solve problems, then we need to move away from exclusive dependence on data models and adopt a more diverse set of tools.