Identifying Possible False Matches in Anonymized Hospital Administrative Data without Patient Identifiers

Identifying Possible False Matches in Anonymized Hospital Administrative Data without Patient Identifiers
复制标题

DOI:
10.1111/1475-6773.12272
复制
发表时间:
2015-08-01
影响因子:
3.4
通讯作者:
Goldstein, Harvey
Goldstein, Harvey
中科院分区:
医学3区
文献类型:
--
作者:
Hagger-Johnson, Gareth;Harron, Katie;Goldstein, Harvey

文献摘要

被引文献

相似文献

目的识别以可能的错误匹配形式出现的数据链接错误,其中两名患者似乎共享相同的唯一识别号。数据来源英国英格兰医院事件统计(HES)。研究设计婴儿出生和再入院数据(2011年4月1日至2012年3月31日; 0 - 1岁)和青少年(2004年4月1日至2011年3月31日;年龄10 - 19岁)。数据收集/提取方法医院记录伪使用设计用于链接属于同一个人的多个记录的算法进行匿名化。六个不可信的临床情况被认为是可能的错误匹配:多胞胎共享HESID,死亡后重新入院,两个出生事件共享HESID,同时在不同的医院住院,婴儿事件编码为分娩,青少年事件编码为births.Principal发现在507,778名婴儿中,可能的错误匹配相对较少(n = 433,0.1%)。最常见的情况(在两家医院同时入院,n = 324)更有可能是缺失数据的婴儿,早产儿和亚洲婴儿。在青少年中,这种情况下(n = 320)是更常见的男性,年轻的患者,混合种族群体,和那些重新承认更frequent.ConclusionsResearchers可以确定临床上难以置信的情况下,受影响的患者,在数据清洗阶段,以减轻可能的连锁错误的影响。
ObjectiveTo identify data linkage errors in the form of possible false matches, where two patients appear to share the same unique identification number.Data SourceHospital Episode Statistics (HES) in England, United Kingdom.Study DesignData on births and re-admissions for infants (April 1, 2011 to March 31, 2012; age 0-1year) and adolescents (April 1, 2004 to March 31, 2011; age 10-19years).Data Collection/Extraction MethodsHospital records pseudo-anonymized using an algorithm designed to link multiple records belonging to the same person. Six implausible clinical scenarios were considered possible false matches: multiple births sharing HESID, re-admission after death, two birth episodes sharing HESID, simultaneous admission at different hospitals, infant episodes coded as deliveries, and adolescent episodes coded as births.Principal FindingsAmong 507,778 infants, possible false matches were relatively rare (n=433, 0.1 percent). The most common scenario (simultaneous admission at two hospitals, n=324) was more likely for infants with missing data, those born preterm, and for Asian infants. Among adolescents, this scenario (n=320) was more common for males, younger patients, the Mixed ethnic group, and those re-admitted more frequently.ConclusionsResearchers can identify clinically implausible scenarios and patients affected, at the data cleaning stage, to mitigate the impact of possible linkage errors.