Demonstrating the robustness of population surveillance data: implications of error rates on demographic and mortality estimates

Demonstrating the robustness of population surveillance data: implications of error rates on demographic and mortality estimates
复制标题

DOI:
10.1186/1471-2288-8-13
复制
发表时间:
2008-03-25
影响因子:
4
通讯作者:
Berhane, Yemane
Berhane, Yemane
中科院分区:
医学3区
文献类型:
--
作者:
Fottrell, Edward;Byass, Peter;Berhane, Yemane

文献摘要

被引文献

相似文献

背景:与任何测量过程一样,常规人口监测操作(例如人口监测站点(DSS)中的操作)可能会出现一定程度的误差。无论使用何种数据采集方法或采用何种质量控制程序,都可能会错过重要事件并出现错误。大型纵向数据集中的随机误差对整体健康和人口概况的影响程度对于 DSS 作为公共卫生研究和临床试验平台的作用具有重要意义。如果要以现实的误差范围和有效性来推断和汇总 DSS 的输出,这些知识也特别重要。 方法:本研究使用埃塞俄比亚布塔吉拉农村卫生项目 (BRHP) DSS 的第一个 10 年数据集,涵盖约 336,000 人年的数据。我们编写了简单的程序,将随机错误和遗漏引入新版本的权威 10 年 Butajira 数据集。选择性别、年龄、死亡、识字率和屋顶材料(贫困指标)等关键参数来引入误差,因为这些参数在人口和健康监测中具有明显的重要性,并且与死亡率存在显着关联。为了本次调查的目的,将原始 10 年数据集定义为“黄金标准”,对每个故意错误的数据集和原始“黄金标准”之间的人口、年龄和性别构成以及死亡率比率的泊松回归模型进行了比较10 年数据。结果:尽管引入了随机误差,但布塔吉拉人口的构成得到了很好的体现,并且基于派生数据集的人口金字塔之间的差异很微妙。即使数据中的随机误差水平相对较高,对既定死亡风险因素的回归分析也基本上不受影响。结论:参数估计和回归分析对大量随机引入的误差的低敏感性表明数据集具有较高的稳健性。总体参数估计对模拟误差的这种明显惯性很大程度上是由于数据集的大小造成的。 DSS 数据中随机误差的容许范围可能超过 20%。虽然这并不是支持低质量数据的论点,但减少在常规 DSS 操作中检测和纠正随机错误所花费的时间和宝贵资源可能是合理的,因为此类程序的回报随着总体准确性的提高而减少。目前花在无休止地纠正 DSS 数据集上的资金和精力也许最好花在增加 DSS 的监测人口规模和地理分布以及分析和传播研究结果上。
Background: As in any measurement process, a certain amount of error may be expected in routine population surveillance operations such as those in demographic surveillance sites (DSSs). Vital events are likely to be missed and errors made no matter what method of data capture is used or what quality control procedures are in place. The extent to which random errors in large, longitudinal datasets affect overall health and demographic profiles has important implications for the role of DSSs as platforms for public health research and clinical trials. Such knowledge is also of particular importance if the outputs of DSSs are to be extrapolated and aggregated with realistic margins of error and validity.Methods: This study uses the first 10-year dataset from the Butajira Rural Health Project (BRHP) DSS, Ethiopia, covering approximately 336,000 person-years of data. Simple programmes were written to introduce random errors and omissions into new versions of the definitive 10-year Butajira dataset. Key parameters of sex, age, death, literacy and roof material (an indicator of poverty) were selected for the introduction of errors based on their obvious importance in demographic and health surveillance and their established significant associations with mortality.Defining the original 10-year dataset as the 'gold standard' for the purposes of this investigation, population, age and sex compositions and Poisson regression models of mortality rate ratios were compared between each of the intentionally erroneous datasets and the original 'gold standard' 10-year data.Results: The composition of the Butajira population was well represented despite introducing random errors, and differences between population pyramids based on the derived datasets were subtle. Regression analyses of well-established mortality risk factors were largely unaffected even by relatively high levels of random errors in the data.Conclusion: The low sensitivity of parameter estimates and regression analyses to significant amounts of randomly introduced errors indicates a high level of robustness of the dataset. This apparent inertia of population parameter estimates to simulated errors is largely due to the size of the dataset. Tolerable margins of random error in DSS data may exceed 20%. While this is not an argument in favour of poor quality data, reducing the time and valuable resources spent on detecting and correcting random errors in routine DSS operations may be justifiable as the returns from such procedures diminish with increasing overall accuracy. The money and effort currently spent on endlessly correcting DSS datasets would perhaps be better spent on increasing the surveillance population size and geographic spread of DSSs and analysing and disseminating research findings.