The Data Science Process: One Culture
The Data Science Process: One Culture
复制标题
数据科学过程:一种文化
DOI:
10.1080/01621459.2020.1762615
复制
发表时间:
2020
影响因子:
3.7
通讯作者:
Barter, Rebecca
中科院分区:
文献类型:
--
作者:
Yu, Bin;Barter, Rebecca
We would like to congratulate Professor Brad Efron for the well-deserved honor of winning the 2019 International Statistics Prize, and for his thought-provoking paper discussing the diverging paths of so-called traditional statistical regression and pure prediction approaches. Professor Efron has made vast contributions to the field of statistics spanning several decades, and we enjoyed reading his take on the current state of the fields of statistics and machine learning. In this discussion, we provide a portrayal of our own perspectives on the intersection of statistics and machine learning methods, and compare them with Professor Efron’s. While we came from very similar beginnings as Professor Efron, rather than arriving at his view of a “discrepancy” or a “radical difference” between “traditional statistical regression methods” and “pure prediction” methods, we arrived at a surprisingly contrasting view that unifies these two approaches. At the very end of his paper, Professor Efron seems “hopeful” about the “reunification” of these two paradigms. However, in his view, efforts toward this reunification are “just underway,” with the primary impediments being a lack of theory. In our contrasting view, we are much further along the path of reunification, with the theoretical underpinnings being less critical than—but still important, and following on from—empirical evidence in today’s realityrooted era.While our views are quite different, we certainly agree with many of the points that Professor Efron makes in his paper. For instance, Professor Efron appears weary of relying solely on random splits when implementing the train/test paradigm that is pervasive in the world of pure prediction algorithms. Specifically, he points out that there are instances where instead of providing an honest estimation of the error, a random split will be far too optimistic, as in the case of time-dependent data. He argues that a more realistic split in such scenarios is not random, but rather uses an early/late split where the early data are used to train and the later data are used to test. Indeed, a random split in many cases ensures that the training and test sets are identically distributed, but fails to ensure that they are independent (ie, a random split does not resemble two independent samples taken from the same population). Such a random split certainly does not resemble the relationship between the data used to train an algorithm and the future data that it will be used on (future data are, at best, an independent sample from the same population, but are often collected from a different source
DOI:
10.1073/pnas.1711236115
发表时间:
2018-02-20
影响因子:
11.1
作者:
Basu S;Kumbier K;Brown JB;Yu B
通讯作者:
Yu B
DOI:
10.1073/pnas.1901326117
发表时间:
2020-02-25
影响因子:
11.1
作者:
Yu, Bin;Kumbier, Karl
通讯作者:
Kumbier, Karl