STATISTICAL PARADISES AND PARADOXES IN BIG DATA (I): LAW OF LARGE POPULATIONS, BIG DATA PARADOX, AND THE 2016 US PRESIDENTIAL ELECTION

STATISTICAL PARADISES AND PARADOXES IN BIG DATA (I): LAW OF LARGE POPULATIONS, BIG DATA PARADOX, AND THE 2016 US PRESIDENTIAL ELECTION
复制标题

DOI:
10.1214/18-aoas1161sf
复制
发表时间:
2018-06-01
影响因子:
1.8
通讯作者:
Meng, Xiao-Li
Meng, Xiao-Li
中科院分区:
数学4区
文献类型:
--
作者:
Meng, Xiao-Li

文献摘要

被引文献

相似文献

统计学家越来越多地提出发人深省甚至自相矛盾的问题,挑战我们进入大数据创造的统计天堂的资格。通过制定数据质量的衡量标准,本文提出了一个框架来解决这样一个问题:“我应该更信任哪一个:一个1%的调查,60%的答复率或自我报告的行政数据集覆盖80%的人口?“一个5元欧拉公式式恒等式表明,对于任何大小为n的数据集,无论概率与否,样本平均值(X)对条(n)和总体平均值(X)对条(N)之间的差异是三项的乘积:(1)数据质量度量,rho(R,X),X-j和响应/记录指标R-j之间的相关性;(2)数据量度量,root(N-n)/n,其中N是总体大小;以及(3)问题难度度量,sigma(X),X的标准差。这种分解提供了多方面的见解:(I)概率抽样通过将rho(R,X)控制在N-1/2的水平来确保高数据质量;(II)当我们失去这种控制时,N的影响不再被rho(R,X)抵消,导致大总体定律(LLP),即我们的估计误差,相对于基准率1/root n,随着root N而增加;以及(III)这种大数据的“巨大性”(对于总体推断)应该用相对大小f = n/N来衡量,而不是绝对大小n;(四)在结合数据源进行人口推断时,那些相对较小但质量较高的公司应该得到比其规模所建议的更多的权重。2016年美国总统大选的一个统计数据表明,自我报告投票给唐纳德·特朗普的rho(R,X)近似为-0.005。由于LLP,这种看似微小的数据缺陷相关性意味着来自1%的美国合格选民的自我报告的特朗普投票偏好的简单样本比例,即n近似于2,300,000,与来自大小为n近似于400的真正简单随机样本的相应样本比例具有相同的均方误差,样本量减少了99.98%(因此我们的置信度也降低了)。CCES的数据生动地展示了LLP:平均而言,该州的选民人数越多,特朗普的实际投票份额就越远离基于样本比例的通常95%置信区间。这应该提醒我们,在不考虑数据质量的情况下,大数据的人口推断受到大数据悖论的影响:数据越多,我们越肯定会欺骗自己。
Statisticians are increasingly posed with thought-provoking and even paradoxical questions, challenging our qualifications for entering the statistical paradises created by Big Data. By developing measures for data quality, this article suggests a framework to address such a question: "Which one should I trust more: a 1% survey with 60% response rate or a self-reported administrative dataset covering 80% of the population?" A 5-element Eulerformula-like identity shows that for any dataset of size n, probabilistic or not, the difference between the sample average (X) over bar (n) and the population average (X) over bar (N) is the product of three terms: (1) a data quality measure, rho(R, X), the correlation between X-j and the response/recording indicator R-j; (2) a data quantity measure, root(N - n)/n, where N is the population size; and (3) a problem difficulty measure, sigma(X), the standard deviation of X. This decomposition provides multiple insights: (I) Probabilistic sampling ensures high data quality by controlling rho(R, X) at the level of N-1/2; (II) When we lose this control, the impact of N is no longer canceled by rho(R, X), leading to a Law of Large Populations (LLP), that is, our estimation error, relative to the benchmarking rate 1/root n, increases with root N; and (III) the "bigness" of such Big Data (for population inferences) should be measured by the relative size f = n/N, not the absolute size n; (IV) When combining data sources for population inferences, those relatively tiny but higher quality ones should be given far more weights than suggested by their sizes.Estimates obtained from the Cooperative Congressional Election Study (CCES) of the 2016 US presidential election suggest a rho(R, X) approximate to -0.005 for self-reporting to vote for Donald Trump. Because of LLP, this seemingly minuscule data defect correlation implies that the simple sample proportion of the self-reported voting preference for Trump from 1% of the US eligible voters, that is, n approximate to 2,300,000, has the same mean squared error as the corresponding sample proportion from a genuine simple random sample of size n approximate to 400, a 99.98% reduction of sample size (and hence our confidence). The CCES data demonstrate LLP vividly: on average, the larger the state's voter populations, the further away the actual Trump vote shares from the usual 95% confidence intervals based on the sample proportions. This should remind us that, without taking data quality into account, population inferences with Big Data are subject to a Big Data Paradox: the more the data, the surer we fool ourselves.