The Data Science Process: One Culture

The Data Science Process: One Culture
复制标题

数据科学过程:一种文化

DOI:
10.1080/01621459.2020.1762615
复制
发表时间:
2020
影响因子:
3.7
通讯作者:
Barter, Rebecca
Barter, Rebecca
中科院分区:
数学1区
文献类型:
--
作者:
Yu, Bin;Barter, Rebecca

文献摘要

参考文献

被引文献

相似文献

我们要祝贺布拉德埃夫隆教授当之无愧的荣誉赢得2019年国际统计奖,并为他发人深省的论文讨论所谓的传统统计回归和纯预测方法的不同路径。埃夫隆教授在几十年的时间里为统计学领域做出了巨大的贡献,我们很喜欢阅读他对统计学和机器学习领域现状的看法。在这次讨论中,我们提供了我们自己对统计和机器学习方法交叉点的观点的描述,并将其与Efron教授的观点进行了比较。虽然我们与埃夫隆教授的起点非常相似,但我们并没有得出他对“传统统计回归方法”和“纯预测”方法之间的“差异”或“根本差异”的看法,而是得出了一个令人惊讶的对比观点,将这两种方法统一起来。在论文的最后,埃夫隆教授似乎对这两种范式的“统一”抱有“希望”。然而,他认为,统一的努力“刚刚开始”,主要障碍是缺乏理论。在我们的对比观点中,我们在统一的道路上走得沿着远,理论基础不如今天现实扎根时代的经验证据那么重要,但仍然很重要,而且是从经验证据中继承下来的。虽然我们的观点截然不同,但我们肯定同意埃夫隆教授在他的论文中提出的许多观点。例如,Efron教授在实现纯预测算法世界中普遍存在的训练/测试范式时,似乎厌倦了仅仅依赖随机分割。具体来说,他指出,在有些情况下,随机分割将过于乐观,而不是提供对误差的诚实估计,就像时间相关数据一样。他认为,在这种情况下,更现实的分割不是随机的,而是使用早期/晚期分割,早期数据用于训练,后期数据用于测试。事实上,在许多情况下,随机分割可以确保训练集和测试集是相同分布的,但不能确保它们是独立的(即随机分割不像从同一个群体中提取的两个独立样本)。这样的随机分割当然不像用于训练算法的数据和将来使用算法的数据之间的关系(将来的数据充其量是来自同一群体的独立样本,但通常是从不同的来源收集的
We would like to congratulate Professor Brad Efron for the well-deserved honor of winning the 2019 International Statistics Prize, and for his thought-provoking paper discussing the diverging paths of so-called traditional statistical regression and pure prediction approaches. Professor Efron has made vast contributions to the field of statistics spanning several decades, and we enjoyed reading his take on the current state of the fields of statistics and machine learning. In this discussion, we provide a portrayal of our own perspectives on the intersection of statistics and machine learning methods, and compare them with Professor Efron’s. While we came from very similar beginnings as Professor Efron, rather than arriving at his view of a “discrepancy” or a “radical difference” between “traditional statistical regression methods” and “pure prediction” methods, we arrived at a surprisingly contrasting view that unifies these two approaches. At the very end of his paper, Professor Efron seems “hopeful” about the “reunification” of these two paradigms. However, in his view, efforts toward this reunification are “just underway,” with the primary impediments being a lack of theory. In our contrasting view, we are much further along the path of reunification, with the theoretical underpinnings being less critical than—but still important, and following on from—empirical evidence in today’s realityrooted era.While our views are quite different, we certainly agree with many of the points that Professor Efron makes in his paper. For instance, Professor Efron appears weary of relying solely on random splits when implementing the train/test paradigm that is pervasive in the world of pure prediction algorithms. Specifically, he points out that there are instances where instead of providing an honest estimation of the error, a random split will be far too optimistic, as in the case of time-dependent data. He argues that a more realistic split in such scenarios is not random, but rather uses an early/late split where the early data are used to train and the later data are used to test. Indeed, a random split in many cases ensures that the training and test sets are identically distributed, but fails to ensure that they are independent (ie, a random split does not resemble two independent samples taken from the same population). Such a random split certainly does not resemble the relationship between the data used to train an algorithm and the future data that it will be used on (future data are, at best, an independent sample from the same population, but are often collected from a different source
DOI: 10.1073/pnas.1711236115
发表时间: 2018-02-20
影响因子: 11.1
作者:
Basu S;Kumbier K;Brown JB;Yu B
通讯作者: Yu B
DOI: 10.1073/pnas.1901326117
发表时间: 2020-02-25
影响因子: 11.1
作者:
Yu, Bin;Kumbier, Karl
通讯作者: Kumbier, Karl