Discussion of Professor Bradley Efron’s Article on “Prediction, Estimation, and Attribution”
Discussion of Professor Bradley Efron’s Article on “Prediction, Estimation, and Attribution”
复制标题
Bradley Efron 教授关于“预测、估计和归因”的文章的讨论
DOI:
10.1111/insr.12415
复制
发表时间:
2020
影响因子:
2
通讯作者:
Zheng, Zheshi
中科院分区:
文献类型:
--
作者:
Xie, Min‐ge;Zheng, Zheshi
By noting the rapid growing trend of “pure prediction algorithms,” Professor Efron compares and bridges the statistics of the 20th Century (estimation and attribution) to that of the current fast growing development of the 21st Century (prediction). The outstanding discussion offers many deep-rooted insights and comments. As did his forward thinking article on Fisher’s influence on modern statistics (Efron 1998), which helped shape many recent developments on statistical inference (including our own work on confidence distribution (Singh, Xie, and Strawderman 2005; Xie and Singh 2013)), this equally inspiring article by Professor Efron will certainly galvanize many contemporary and powerful developments for modern statistics and for the foundations of data science. In this note, we echo and also provide additional support to two important points made by Professor Efron:(1) prediction is “an easier task than either attribution or estimation”;(2) the IID assumption (eg random splitting of training and testing datasets) is crucial in the current developments on predictions, but we also need to do more for the case when the IID assumption is not met. Based on our own research, we provide additional evidence to support these discussions. We discover that prediction has a homeostasis property and works well under the IID setting even if the learning model used is completely wrong. We also highlight the importance of having a good modeling and inference practice: a good learning model with good estimation is important to improve prediction efficiency in the IID case and it becomes essential to maintain validity in the non-IID case. The message remains: we still need to make effort to build a good learning model and estimation algorithm in prediction, even if prediction is an easier task than estimation.From the outset, we would like to point out that it is not a straw-man argument to consider non-IID testing data. On the contrary, such data are prevalent in data science. In addition to those examples provided by Professor Efron that showed “drift,” we can easily imagine non-IID examples in many typical applications. For instance, a predictive algorithm is trained on a database of patient medical records and we would like to predict potential outcomes of a treatment for a new patient with more severe symptoms than what the average patient shows. The new patient with more severe symptoms is not a typical
影响因子:
1.6
作者:
Barber, Rina Foygel;Candes, Emmanuel J.;Tibshirani, Ryan J.
通讯作者:
Tibshirani, Ryan J.
影响因子:
4.5
作者:
Barber, Rina Foygel;Candes, Emmanuel J.;Tibshirani, Ryan J.
通讯作者:
Tibshirani, Ryan J.