Big Data and Large Sample Size: A Cautionary Note on the Potential for Bias

Big Data and Large Sample Size: A Cautionary Note on the Potential for Bias
复制标题

DOI:
10.1111/cts.12178
复制
发表时间:
2014-08-01
影响因子:
3.9
通讯作者:
Glasgow, Russell E.
Glasgow, Russell E.
中科院分区:
医学3区
文献类型:
--
作者:
Kaplan, Robert M.;Chambers, David A.;Glasgow, Russell E.

文献摘要

被引文献

相似文献

一些评论指出,大型研究比小型研究更可靠,人们对综合了数千人和/或不同数据来源的信息的“大数据”的分析越来越感兴趣。我们考虑了大数据时代可能出现的各种偏差,包括抽样错误、测量错误、多次比较错误、聚合错误以及与系统排除信息相关的错误。利用流行病学、卫生服务研究、健康决定因素研究和临床试验的例子,我们得出结论,有必要更加谨慎,以确保大样本量不会导致大的推断错误。尽管大研究有优势,但大样本量可能会放大抽样或研究设计引起的误差相关的偏差。
A number of commentaries have suggested that large studies are more reliable than smaller studies and there is a growing interest in the analysis of "big data" that integrates information from many thousands of persons and/or different data sources. We consider a variety of biases that are likely in the era of big data, including sampling error, measurement error, multiple comparisons errors, aggregation error, and errors associated with the systematic exclusion of information. Using examples from epidemiology, health services research, studies on determinants of health, and clinical trials, we conclude that it is necessary to exercise greater caution to be sure that big sample size does not lead to big inferential errors. Despite the advantages of big studies, large sample size can magnify the bias associated with error resulting from sampling or study design.