Unrepresentative big surveys significantly overestimated US vaccine uptake

Unrepresentative big surveys significantly overestimated US vaccine uptake
复制标题

DOI:
10.1038/s41586-021-04198-4
复制
发表时间:
2021-12-08
期刊:
影响因子:
64.8
通讯作者:
Flaxman, Seth
Flaxman, Seth
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Bradley, Valerie C.;Kuriwaki, Shiro;Flaxman, Seth

文献摘要

被引文献

相似文献

调查是了解公众舆论和行为的重要工具,其准确性取决于通过尽量减少各种来源的偏见来保持其目标人口的统计代表性。增加数据大小缩小了置信区间,但放大了调查偏差的影响:大数据Partial的一个实例(1)。在此,我们从两项大型调查中对2021年1月9日至5月19日美国成年人首次接种COVID-19疫苗的估计值证明了这一悖论:Delphi-Facebook(2,3)(每周约250,000人回复)和Census Household Pulse(4)(每两周约75,000人回复)。2021年5月,与美国疾病控制和预防中心于2021年5月26日发布的追溯更新基准相比,Delphi-Facebook高估了17个百分点(14-20个百分点,基准不精确度为5%),而Census Household Pulse则高估了14个百分点(11-17个百分点,基准不精确度为5%)。此外,他们的样本量很大,导致不正确估计的误差很小。相比之下,Axios-Ipsos在线小组(5)根据调查研究最佳实践(6)每周约有1,000份回复,提供了可靠的估计和不确定性量化。我们使用最近的分析框架(1)分解观察到的误差,以解释三次调查中的不准确性。然后,我们分析了疫苗犹豫和意愿的影响。我们展示了一个对25万名受访者的调查如何产生一个人口平均值的估计,这个估计并不比一个简单的随机样本的估计更准确。我们的中心信息是,数据质量比数据数量更重要,用后者来补偿前者是一个数学上可以证明的失败主张。
Surveys are a crucial tool for understanding public opinion and behaviour, and their accuracy depends on maintaining statistical representativeness of their target populations by minimizing biases from all sources. Increasing data size shrinks confidence intervals but magnifies the effect of survey bias: an instance of the Big Data Paradox(1). Here we demonstrate this paradox in estimates of first-dose COVID-19 vaccine uptake in US adults from 9 January to 19 May 2021 from two large surveys: Delphi-Facebook(2,3) (about 250,000 responses per week) and Census Household Pulse(4) (about 75,000 every two weeks). In May 2021, Delphi-Facebook overestimated uptake by17 percentage points (14-20 percentage points with 5% benchmark imprecision) and Census Household Pulse by 14 (11-17 percentage points with 5% benchmark imprecision), compared to a retroactively updated benchmark the Centers for Disease Control and Prevention published on 26 May 2021. Moreover, their large sample sizes led to miniscule margins of error on the incorrect estimates. By contrast, an Axios-Ipsos online panel(5) with about 1,000 responses per week following survey research best practices(6) provided reliable estimates and uncertainty quantification. We decompose observed error using a recent analytic framework(1) to explain the inaccuracy in the three surveys. We then analyse the implications for vaccine hesitancy and willingness. We show how a survey of 250,000 respondents can produce an estimate of the population mean that is no more accurate than an estimate from a simple random sample of size 10. Our central message is that data quality matters more than data quantity, and that compensating the former with the latter is a mathematically provable losing proposition.