Big Data and the danger of being precisely inaccurate

Big Data and the danger of being precisely inaccurate
复制标题

大数据和精确不准确的危险

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
8.5
通讯作者:
Richard McFarland
Richard McFarland
中科院分区:
法学1区
文献类型:
--
作者:
Daniel A. McFarland;Richard McFarland

文献摘要

被引文献

相似文献

社会科学家和数据分析师越来越多地在其分析中使用大数据。轻松地满足大多数统计程序的样本要求,他们在专注于采用传统统计方法时使分析师具有错误的安全感解释当今大数据对大数据进行的大多数分析导致“确切的不准确”结果,这些结果隐藏了数据中的偏见,但由于数据大小在大型数据集执行任何分析的结果的增强,因此很容易被忽略。我们建议采用简单的数据分割技术来控制观察数据偏见的一些主要组成部分。
Social scientists and data analysts are increasingly making use of Big Data in their analyses. These data sets are often “found data” arising from purely observational sources rather than data derived under strict rules of a statistically designed experiment. However, since these large data sets easily meet the sample size requirements of most statistical procedures, they give analysts a false sense of security as they proceed to focus on employing traditional statistical methods. We explain how most analyses performed on Big Data today lead to “precisely inaccurate” results that hide biases in the data but are easily overlooked due to the enhanced significance of the results created by the data size. Before any analyses are performed on large data sets, we recommend employing a simple data segmentation technique to control for some major components of observational data biases. These segments will help to improve the accuracy of the results.