PREDICTIVE MODELING WITH BIG DATA: Is Bigger Really Better?

PREDICTIVE MODELING WITH BIG DATA: Is Bigger Really Better?
复制标题

DOI:
10.1089/big.2013.0037
复制
发表时间:
2013-12-01
期刊:
影响因子:
4.6
通讯作者:
Provost, Foster
Provost, Foster
中科院分区:
计算机科学4区
文献类型:
--
作者:
de Fortuny, Enric Junque;Martens, David;Provost, Foster

文献摘要

被引文献

相似文献

随着“大数据”的收集和处理日益广泛,人们自然有兴趣使用这些数据资产来改善决策。使用数据来改善决策的最佳理解方式之一是通过预测分析。一个重要的开放性问题是:在多大程度上更大的数据实际上会导致更好的预测模型?在这篇文章中,我们以经验证明,当预测模型是从稀疏的细粒度数据(如低层次的人类行为数据)中构建时,即使在非常大的规模上,我们仍然可以看到预测性能的边际增长。实证结果是基于从9个不同的预测建模应用程序,从书评到银行交易的数据。这项研究清楚地表明,更大的数据确实可以成为预测分析的更有价值的资产。这意味着拥有更大数据资产的机构-加上利用它们的技能-可能会获得比没有这种访问或技能的机构更大的竞争优势。此外,研究结果表明,对于能够访问此类细粒度数据的公司来说,在关键预测任务的背景下,收集更多的数据实例和更多可能的数据特征是值得的。作为一个额外的贡献,我们介绍了多变量Bernoulli朴素贝叶斯算法,可以扩展到大量的,稀疏的数据的实现。
With the increasingly widespread collection and processing of "big data," there is natural interest in using these data assets to improve decision making. One of the best understood ways to use data to improve decision making is via predictive analytics. An important, open question is: to what extent do larger data actually lead to better predictive models? In this article we empirically demonstrate that when predictive models are built from sparse, fine-grained data-such as data on low-level human behavior-we continue to see marginal increases in predictive performance even to very large scale. The empirical results are based on data drawn from nine different predictive modeling applications, from book reviews to banking transactions. This study provides a clear illustration that larger data indeed can be more valuable assets for predictive analytics. This implies that institutions with larger data assets-plus the skill to take advantage of them-potentially can obtain substantial competitive advantage over institutions without such access or skill. Moreover, the results suggest that it is worthwhile for companies with access to such fine-grained data, in the context of a key predictive task, to gather both more data instances and more possible data features. As an additional contribution, we introduce an implementation of the multivariate Bernoulli Naive Bayes algorithm that can scale to massive, sparse data.