Methods for dealing with discrepant records in linked population health datasets: a cross-sectional study

Methods for dealing with discrepant records in linked population health datasets: a cross-sectional study
复制标题

DOI:
10.1186/1472-6963-7-12
复制
发表时间:
2007-01-30
影响因子:
2.8
通讯作者:
Ford, Jane B.
Ford, Jane B.
中科院分区:
医学3区
文献类型:
--
作者:
Roberts, Christine L.;Algert, Charles S.;Ford, Jane B.

文献摘要

被引文献

相似文献

背景:相关的人口健康数据越来越多地用于流行病学研究。如果在多个数据集上报告数据项,数据链接可以减少与许多人口健康数据集相关的不充分确定。然而,这增加了来自不同数据集的不同病例报告的可能性。方法:我们研究了四种方法对来自不同人群健康数据集的差异报告进行分类,对妊娠期高血压疾病的估计患病率和已知危险因素的调整优势比(aOR)的影响。数据来自2000年至2002年在新南威尔士州一家医院(澳大利亚)分娩的妇女的关联、有效的分娩和住院数据。结果:在250173名有关联数据的女性中,238412名(95.3%)女性对高血压的发生完全一致,1577名(0.6%)女性不完全一致;9369名(3.7%)患者仅在一个数据集中报告了高血压(漏报),815名(0.3%)患者的高血压类型相互冲突。仅使用出生和出院数据之间的完美一致性导致最低的患病率(0.3%慢性,5.1%妊娠高血压),而包括所有报告导致最高的患病率(1.1%慢性,8.7%妊娠高血压)。较高的流行率总体上与国际报告一致。相反,对于已知的危险因素,完全一致给出了最高的aOR(95%置信区间):母亲年龄>= 40岁的慢性高血压风险为4.0(2.9,5.3),多胎妊娠高血压风险为2.8(2.5,3.2)。结论:对差异病例报告的分类方法应根据研究问题的不同而有所不同;所有报告都应作为计算患病率估计范围的一部分,但完美匹配可能最适合于风险因素分析。这些发现可能适用于任何专门卫生服务数据集与包括诊断或程序信息的人口数据的联系。
Background: Linked population health data are increasingly used in epidemiological studies. If data items are reported on more than one dataset, data linkage can reduce the under-ascertainment associated with many population health datasets. However, this raises the possibility of discrepant case reports from different datasets.Methods: We examined the effect of four methods of classifying discrepant reports from different population health datasets on the estimated prevalence of hypertensive disorders of pregnancy and on the adjusted odds ratios (aOR) for known risk factors. Data were obtained from linked, validated, birth and hospital data for women who gave birth in a New South Wales hospital ( Australia) 2000 - 2002.Results: Among 250173 women with linked data, 238412 (95.3%) women had perfect agreement on the occurrence of hypertension, 1577 (0.6%) had imperfect agreement; 9369 (3.7%) had hypertension reported in only one dataset (under-reporting) and 815 (0.3%) had conflicting types of hypertension. Using only perfect agreement between birth and discharge data resulted in the lowest prevalence rates ( 0.3% chronic, 5.1% pregnancy hypertension), while including all reports resulted in the highest prevalence rates (1.1 % chronic, 8.7% pregnancy hypertension). The higher prevalence rates were generally consistent with international reports. In contrast, perfect agreement gave the highest aOR (95% confidence interval) for known risk factors: risk of chronic hypertension for maternal age >= 40 years was 4.0 (2.9, 5.3) and the risk of pregnancy hypertension for multiple birth was 2.8 (2.5, 3.2).Conclusion: The method chosen for classifying discrepant case reports should vary depending on the study question; all reports should be used as part of calculating the range of prevalence estimates, but perfect matches may be best suited to risk factor analyses. These findings are likely to be applicable to the linkage of any specialised health services datasets to population data that include information on diagnoses or procedures.