Data linkage errors in hospital administrative data when applying a pseudonymisation algorithm to paediatric intensive care records

Data linkage errors in hospital administrative data when applying a pseudonymisation algorithm to paediatric intensive care records
复制标题

DOI:
10.1136/bmjopen-2015-008118
复制
发表时间:
2015-01-01
期刊:
影响因子:
2.9
通讯作者:
Parslow, Roger C.
Parslow, Roger C.
中科院分区:
医学3区
文献类型:
--
作者:
Hagger-Johnson, Gareth;Harron, Katie;Parslow, Roger C.

文献摘要

被引文献

相似文献

目的:我们的目的是通过在儿科重症监护记录的国家登记处测试HESID伪匿名算法来估计医院事件统计(HES)中的数据链接错误率。设置:儿科重症监护审计网络(PICANet)数据库,涵盖英格兰,苏格兰和威尔士的33个儿科重症监护病房。参与者:数据来自2004年1月1日至2014年2月21日期间入院的0-19岁婴儿和青少年。主要和次要结局指标:PICANet的入院记录被归类为匹配在将HESID算法应用于PICANet记录之后,可以识别不匹配(记录属于再次入院的同一患者)或不匹配(记录属于不同患者)。通过比较HESID算法与参考标准PICANetID. Effectoflinkingerrors对再入院率的影响,计算假匹配和错配率。结果:在166406例入院病例中,88596例为真匹配(同一患者再次入院)。HESID自动化算法产生了很少的错误匹配(n= 176/77 810; 0.2%),但较大比例的漏匹配(n= 3609/88 596; 4.1%)。由于连锁错误,真实再入院率被低估了3.8%。来自亚洲/黑人/其他种族(vs白色)的年轻男性患者更有可能出现假匹配。错过的比赛是更常见的年轻患者,亚洲/黑人/其他种族群体(与白色)和患者的记录有缺失data.Conclusions:确定性算法用于连接所有事件的医院护理在英格兰同一患者有很高的错过匹配率,低估了真正的再入院率,并会产生偏见的分析。为了减少链接错误,伪匿名化算法需要根据高质量的参考标准进行验证。数据“在源”的假名化本身并不能解决患者标识符中的错误以及这些错误对数据链接的影响。
Objectives: Our aim was to estimate the rate of data linkage error in Hospital Episode Statistics (HES) by testing the HESID pseudoanonymisation algorithm against a reference standard, in a national registry of paediatric intensive care records.Setting: The Paediatric Intensive Care Audit Network (PICANet) database, covering 33 paediatric intensive care units in England, Scotland and Wales.Participants: Data from infants and young people aged 0-19 years admitted between 1 January 2004 and 21 February 2014.Primary and secondary outcome measures: PICANet admission records were classified as matches (records belonging to the same patient who had been readmitted) or non-matches (records belonging to different patients) after applying the HESID algorithm to PICANet records. False-match and missed-match rates were calculated by comparing results of the HESID algorithm with the reference standard PICANet ID. The effect of linkage errors on readmission rate was evaluated.Results: Of 166 406 admissions, 88 596 were true matches (where the same patient had been readmitted). The HESID pseudonymisation algorithm produced few false matches (n= 176/77 810; 0.2%) but a larger proportion of missed matches (n= 3609/88 596; 4.1%). The true readmission rate was underestimated by 3.8% due to linkage errors. Patients who were younger, male, from Asian/Black/Other ethnic groups (vs White) were more likely to experience a false match. Missed matches were more common for younger patients, for Asian/Black/Other ethnic groups (vs White) and for patients whose records had missing data.Conclusions: The deterministic algorithm used to link all episodes of hospital care for the same patient in England has a high missed match rate which underestimates the true readmission rate and will produce biased analyses. To reduce linkage error, pseudoanonymisation algorithms need to be validated against good quality reference standards. Pseudonymisation of data 'at source' does not itself address errors in patient identifiers and the impact these errors have on data linkage.