Accuracy of commercial geocoding: assessment and implications.

Accuracy of commercial geocoding: assessment and implications.
复制标题

DOI:
10.1186/1742-5573-3-8
复制
发表时间:
2006-07-20
期刊:
Epidemiologic perspectives & innovations : EP+I
影响因子:
--
通讯作者:
Heiss, Gerardo
Heiss, Gerardo
中科院分区:
其他
文献类型:
--
作者:
Whitsel, Eric A;Quibrera, P Miguel;Heiss, Gerardo

文献摘要

被引文献

相似文献

背景:已发表的地理编码准确性研究通常集中于单个地理区域、地址源或供应商,不调整地址特征的准确性测量,也不检查不准确性对暴露测量的影响。我们在妇女健康倡议的一项辅助研究《WHI 中心律失常发生的环境流行病学》中解决了这些问题。结果:美国 49 个州 (n = 3,615) 的已确定坐标的地址由四家供应商 (A-D) 进行了地理编码。供应商之间在地址匹配率(98%;82%;81%;30%)、已建立的人口普查区与供应商分配的人口普查区之间的一致性(85%;88%;87%;98%)以及已建立的坐标与供应商分配的坐标之间的距离(平均 rho [米]:1809;748;704;228)方面存在显着差异。在街道匹配、完整、邮政编码、未经编辑的城市地址以及采用 1983 年北美基准面或 1984 年世界大地测量系统坐标的地址中,平均 rho 最低。在仅限于具有最低可接受匹配率 (A-C) 的供应商并针对地址特征、地址内相关性和 rho 的供应商间异方差性进行调整的混合模型中,街道类型匹配的平均 rho 差异很小(280;268;275),即对于大多数应用程序来说,可能会导致依赖于它们的结果产生大致相同的偏差。相比之下,在某些供应商对比中,质心类型匹配之间的差异很大,但在其他供应商对比中则不然 (5497; 4303; 4210) p(交互) < 10(-4),即在许多应用中更有可能使结果产生不同的偏差。供应商 A 与 C 的地址匹配调整后赔率较高(赔率 = 66,95% 置信区间:47、93),但 B 与 C 相比则不然(OR = 1.1,95% CI:0.9、1.3)。供应商 A 与 C(OR = 1.0,95% CI:0.9、1.2)或 B 与 C(OR = 1.1,95% CI:0.9、1.3)的人口普查区一致性并不更高。相关暴露测量(到最近高速公路的距离)的错误分类随着平均 rho 的增加而增加,并且在没有混杂因素的情况下,该距离的非差异性错误分类使其与冠心病死亡率的假设关联偏向零。结论:地理编码错误取决于用于评估它的措施、地址特征和供应商。供应商选择提出了数据丢失的可能性和估计空间定义属性的错误之间的权衡。需要明智的选择来控制权衡并调整其效果的分析。
BACKGROUND: Published studies of geocoding accuracy often focus on a single geographic area, address source or vendor, do not adjust accuracy measures for address characteristics, and do not examine effects of inaccuracy on exposure measures. We addressed these issues in a Women's Health Initiative ancillary study, the Environmental Epidemiology of Arrhythmogenesis in WHI.RESULTS: Addresses in 49 U.S. states (n = 3,615) with established coordinates were geocoded by four vendors (A-D). There were important differences among vendors in address match rate (98%; 82%; 81%; 30%), concordance between established and vendor-assigned census tracts (85%; 88%; 87%; 98%) and distance between established and vendor-assigned coordinates (mean rho [meters]: 1809; 748; 704; 228). Mean rho was lowest among street-matched, complete, zip-coded, unedited and urban addresses, and addresses with North American Datum of 1983 or World Geodetic System of 1984 coordinates. In mixed models restricted to vendors with minimally acceptable match rates (A-C) and adjusted for address characteristics, within-address correlation, and among-vendor heteroscedasticity of rho, differences in mean rho were small for street-type matches (280; 268; 275), i.e. likely to bias results relying on them about equally for most applications. In contrast, differences between centroid-type matches were substantial in some vendor contrasts, but not others (5497; 4303; 4210) p(interaction) < 10(-4), i.e. more likely to bias results differently in many applications. The adjusted odds of an address match was higher for vendor A versus C (odds ratio = 66, 95% confidence interval: 47, 93), but not B versus C (OR = 1.1, 95% CI: 0.9, 1.3). That of census tract concordance was no higher for vendor A versus C (OR = 1.0, 95% CI: 0.9, 1.2) or B versus C (OR = 1.1, 95% CI: 0.9, 1.3). Misclassification of a related exposure measure--distance to the nearest highway--increased with mean rho and in the absence of confounding, non-differential misclassification of this distance biased its hypothetical association with coronary heart disease mortality toward the null.CONCLUSION: Geocoding error depends on measures used to evaluate it, address characteristics and vendor. Vendor selection presents a trade-off between potential for missing data and error in estimating spatially defined attributes. Informed selection is needed to control the trade-off and adjust analyses for its effects.