Spatial aggregation choice in the era of digital and administrative surveillance data.

Spatial aggregation choice in the era of digital and administrative surveillance data.
复制标题

DOI:
10.1371/journal.pdig.0000039
复制
发表时间:
2022-06
期刊:
PLOS digital health
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

相似文献

传统的疾病监测越来越多地得到来自非传统来源的数据的补充,如医疗索赔,电子健康记录和参与性综合征数据平台。由于非传统数据往往是在个人层面收集的,而且是从人口中抽取的方便样本,因此必须选择将这些数据汇总起来进行流行病学推断。我们的研究旨在了解空间聚集选择对我们理解疾病传播的影响,以美国流感样疾病为例。使用2002年至2009年美国医疗索赔数据,我们检查了流行源位置,发病和高峰季节时间,以及流感季节的流行持续时间,以汇总到县和州的规模。我们还比较了空间自相关性,并测试了疾病负担的发病和峰值措施之间的空间聚集差异的相对大小。我们发现,在比较县级和州级数据时,推断的流行源位置和估计的流感季节发病和高峰之间存在差异。在更广阔的地理范围内检测到的空间自相关在高峰季节相比,早期的流感季节,有更大的空间聚集差异,以及在早期的季节措施。在美国流感季节的早期,流行病学推断对空间尺度更敏感,当时流行病的时间,强度和地理传播具有更大的异质性。非传统疾病监测的用户应仔细考虑如何从更精细的数据中提取准确的疾病信号,以便在疾病暴发时及早使用。行政健康记录、Twitter等社交媒体流以及流感网络等参与式监测系统越来越多地用于传染病监测,但通常在地理上进行汇总,以保护数据隐私和机密性。我们探讨了非传统疾病数据源的空间聚集中的任意选择如何影响疾病负担的估计和对疫情的流行病学理解。使用流感样疾病的医疗索赔数据库作为我们的案例研究,我们发现,有很大的变化,在流感季节的时间和幅度跨空间尺度,由于空间聚集可能导致误导性的流行病学数量的估计。特别是,我们发现,流行病学的推断是更敏感的空间尺度在美国流感季节早期,当有更大的异质性,在时间,强度和地理传播的流行病。非传统疾病监测在报告速度和数量方面可能具有明显的优势,但在为空间流行病学分析汇总这些数据时需要谨慎。
Traditional disease surveillance is increasingly being complemented by data from non-traditional sources like medical claims, electronic health records, and participatory syndromic data platforms. As non-traditional data are often collected at the individual-level and are convenience samples from a population, choices must be made on the aggregation of these data for epidemiological inference. Our study seeks to understand the influence of spatial aggregation choice on our understanding of disease spread with a case study of influenza-like illness in the United States. Using U.S. medical claims data from 2002 to 2009, we examined the epidemic source location, onset and peak season timing, and epidemic duration of influenza seasons for data aggregated to the county and state scales. We also compared spatial autocorrelation and tested the relative magnitude of spatial aggregation differences between onset and peak measures of disease burden. We found discrepancies in the inferred epidemic source locations and estimated influenza season onsets and peaks when comparing county and state-level data. Spatial autocorrelation was detected across more expansive geographic ranges during the peak season as compared to the early flu season, and there were greater spatial aggregation differences in early season measures as well. Epidemiological inferences are more sensitive to spatial scale early on during U.S. influenza seasons, when there is greater heterogeneity in timing, intensity, and geographic spread of the epidemics. Users of non-traditional disease surveillance should carefully consider how to extract accurate disease signals from finer-scaled data for early use in disease outbreaks. Administrative health records, social media streams like Twitter, and participatory surveillance systems like Influenzanet are increasingly available for infectious disease surveillance, but are often geographically aggregated to preserve data privacy and confidentiality. We explored how an arbitrary choice in the spatial aggregation of non-traditional disease data sources may influence estimates of disease burden and epidemiological understanding of an outbreak. Using influenza-like illness as measured through a medical claims database as our case study, we find that there is substantial variation in influenza season timing and magnitude across spatial scales due to which spatial aggregation could lead to misleading estimates of epidemiological quantities. In particular, we find that epidemiological inferences are more sensitive to spatial scale early on during U.S. influenza seasons, when there is greater heterogeneity in timing, intensity, and geographic spread of the epidemics. Non-traditional disease surveillance may have distinct advantages in reporting speed and volume, but care is required when aggregating this data for spatial epidemiological analysis.