50 Years of Data Science

50 Years of Data Science
复制标题

DOI:
10.1080/10618600.2017.1384734
复制
发表时间:
2017-01-01
影响因子:
2.4
通讯作者:
Donoho, David
Donoho, David
中科院分区:
数学2区
文献类型:
--
作者:
Donoho, David

文献摘要

被引文献

相似文献

50多年前,约翰·杜克(John Tukey)呼吁对学术统计进行改革。在《数据分析的未来》(The Future of Data Analysis)一书中,他指出了一门尚未被认可的科学的存在,其感兴趣的主题是从数据中学习,或称“数据分析”。十到二十年前,约翰·钱伯斯、杰夫·吴、比尔·克利夫兰和里奥·布雷曼各自独立地再次敦促学术统计超越理论统计的经典领域;钱伯斯呼吁更多地强调数据的准备和呈现,而不是统计建模;Breiman呼吁强调预测而不是推断。克利夫兰和吴甚至建议为这个设想中的领域取一个朗朗上口的名字“数据科学”。最近,越来越多的大学出现了“数据科学”项目,包括加州大学伯克利分校(UC Berkeley)、纽约大学(NYU)、麻省理工学院(MIT),最引人注目的是密歇根大学(University of Michigan)。密歇根大学在2015年9月宣布了一项1亿美元的“数据科学计划”(data science Initiative),旨在招聘35名新教师。这些新课程的教学在课程内容上与传统的统计学课程有很大的重叠;然而,许多学术统计学家认为这些新项目是“文化挪用”。本文回顾了当前“数据科学时刻”的一些因素,包括大众媒体最近对数据科学的评论,以及数据科学如何/是否真的不同于统计学。现在考虑的数据科学领域相当于统计学和机器学习领域的超集,它增加了一些“扩展”到“大数据”的技术。这个被选中的超级群体的动机是商业发展,而不是智力发展。以这种方式选择很可能会错过未来50年真正重要的智力事件。因为所有的科学本身很快就会变成可以挖掘的数据,数据科学即将发生的革命不仅仅是“扩大规模”,而是在科学范围内出现数据分析的科学研究。在未来,我们将能够预测改变数据分析工作流程的提议将如何影响所有科学领域数据分析的有效性,甚至可以逐领域预测影响。根据Tukey、Cleveland、Chambers和Breiman的工作,我提出了一种基于“从数据中学习”的人的活动的数据科学愿景,并描述了一个致力于以循证方式改进这种活动的学术领域。与今天的数据科学计划相比,这个新领域是统计学和机器学习的更好的学术扩展,同时能够适应相同的短期目标。基于2015年9月18日在新泽西州普林斯顿的杜克大学百年纪念研讨会上的一次演讲。
More than 50 years ago, John Tukey called for a reformation of academic statistics. In "The Future of Data Analysis," he pointed to the existence of an as-yet unrecognized science, whose subject of interest was learning from data, or "data analysis." Ten to 20 years ago, John Chambers, Jeff Wu, Bill Cleveland, and Leo Breiman independently once again urged academic statistics to expand its boundaries beyond the classical domain of theoretical statistics; Chambers called for more emphasis on data preparation and presentation rather than statistical modeling; and Breiman called for emphasis on prediction rather than inference. Cleveland and Wu even suggested the catchy name "data science" for this envisioned field. A recent and growing phenomenon has been the emergence of "data science" programs at major universities, including UC Berkeley, NYU, MIT, and most prominently, the University of Michigan, which in September 2015 announced a $100M "Data Science Initiative" that aims to hire 35 new faculty. Teaching in these new programs has significant overlap in curricular subject matter with traditional statistics courses; yet many academic statisticians perceive the new programs as "cultural appropriation." This article reviews some ingredients of the current "data science moment," including recent commentary about data science in the popular media, and about how/whether data science is really different from statistics. The now-contemplated field of data science amounts to a superset of the fields of statistics and machine learning, which adds some technology for "scaling up" to "big data." This chosen superset is motivated by commercial rather than intellectual developments. Choosing in this way is likely to miss out on the really important intellectual event of the next 50 years. Because all of science itself will soon become data that can be mined, the imminent revolution in data science is not about mere "scaling up," but instead the emergence of scientific studies of data analysis science-wide. In the future, we will be able to predict how a proposal to change data analysis workflows would impact the validity of data analysis across all of science, even predicting the impacts field-by-field. Drawing on work by Tukey, Cleveland, Chambers, and Breiman, I present a vision of data science based on the activities of people who are "learning from data," and I describe an academic field dedicated to improving that activity in an evidence-based manner. This new field is a better academic enlargement of statistics and machine learning than today's data science initiatives, while being able to accommodate the same short-term goals. Based on a presentation at the Tukey Centennial Workshop, Princeton, NJ, September 18, 2015.