Big data need big theory too.

Big data need big theory too.
复制标题

DOI:
10.1098/rsta.2016.0153
复制
发表时间:
2016-11-13
期刊:
Philosophical transactions. Series A, Mathematical, physical, and engineering sciences
影响因子:
--
通讯作者:
Highfield RR
Highfield RR
中科院分区:
其他
文献类型:
--
作者:
Coveney PV;Dougherty ER;Highfield RR

文献摘要

参考文献

被引文献

相似文献

当前对大数据,机器学习和数据分析的兴趣产生了宽度的印象,即这种方法能够解决大多数问题,而无需传统的科学探究方法。从科学,医疗保健和网络安全到经济学,社会科学和人类中,可以在几乎所有努力领域中获取数据。学习似乎提供了一个捷径,以揭示原子,分子,中索和宏观的过程之间的任意复杂性的相关性,我们指出了纯粹的大数据方法的弱点,这些弱点未能提供对生物学和医学的特定关注。不管它们的“深度”以及数据驱动的方法(例如人造神经网络)的应用数据不仅需要比大数据狂热者所预期的要多的数据来产生统计上可靠的结果为了建模基础系统的结构特征。和概念知识。生活,医学和医疗保健。 本文是主题问题的一部分,“物理 - 化学与生物学接口的多尺度建模”。
The current interest in big data, machine learning and data analytics has generated the widespread impression that such methods are capable of solving most problems without the need for conventional scientific methods of inquiry. Interest in these methods is intensifying, accelerated by the ease with which digitized data can be acquired in virtually all fields of endeavour, from science, healthcare and cybersecurity to economics, social sciences and the humanities. In multiscale modelling, machine learning appears to provide a shortcut to reveal correlations of arbitrary complexity between processes at the atomic, molecular, meso- and macroscales. Here, we point out the weaknesses of pure big data approaches with particular focus on biology and medicine, which fail to provide conceptual accounts for the processes to which they are applied. No matter their ‘depth’ and the sophistication of data-driven methods, such as artificial neural nets, in the end they merely fit curves to existing data. Not only do these methods invariably require far larger quantities of data than anticipated by big data aficionados in order to produce statistically reliable results, but they can also fail in circumstances beyond the range of the data used to train them because they are not designed to model the structural characteristics of the underlying system. We argue that it is vital to use theory as a guide to experimental design for maximal efficiency of data collection and to produce reliable predictive models and conceptual knowledge. Rather than continuing to fund, pursue and promote ‘blind’ big data projects with massive budgets, we call for more funding to be allocated to the elucidation of the multiscale and stochastic processes controlling the behaviour of complex systems, including those of life, medicine and healthcare. This article is part of the themed issue ‘Multiscale modelling at the physics–chemistry–biology interface’.
DOI: 10.1371/journal.pmed.0020124
发表时间: 2005-08-01
期刊: PLOS MEDICINE
影响因子: 15.8
作者:
Ioannidis, JPA
通讯作者: Ioannidis, JPA
DOI: 10.7554/elife.00747
发表时间: 2013-06-25
期刊: eLife
影响因子: 7.7
作者:
Bozic I;Reiter JG;Allen B;Antal T;Chatterjee K;Shah P;Moon YS;Yaqubie A;Kelly N;Le DT;Lipson EJ;Chapman PB;Diaz LA Jr;Vogelstein B;Nowak MA
通讯作者: Nowak MA
DOI: 10.1039/c6cp02349e
发表时间: 2016-01-01
影响因子: 3.3
作者:
Coveney, Peter V.;Wan, Shunzhou
通讯作者: Wan, Shunzhou
DOI: 10.2174/1573409912666160120151627
发表时间: 2016-01-01
影响因子: 1.7
作者:
Balasubramanian, Krishnan;Basak, Subhash C.
通讯作者: Basak, Subhash C.
DOI: 10.1093/bioinformatics/btt205
发表时间: 2013-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Chowdhury SA;Shackney SE;Heselmeyer-Haddad K;Ried T;Schäffer AA;Schwartz R
通讯作者: Schwartz R