Combining knowledge- and data-driven methods for de-identification of clinical narratives.

Combining knowledge- and data-driven methods for de-identification of clinical narratives.
复制标题

DOI:
10.1016/j.jbi.2015.06.029
复制
发表时间:
2015-12
影响因子:
4.5
通讯作者:
Nenadic G
Nenadic G
中科院分区:
医学3区
文献类型:
--
作者:
Dehghan A;Kovacevic A;Karystianis G;Keane JA;Nenadic G

文献摘要

被引文献

相似文献

最近承诺从电子健康记录中大规模访问非结构化临床数据,这重新激发了人们对临床笔记自动识别的兴趣,其中包括识别提及受保护的健康信息(PHI)。我们描述了作为i2b2/UTHealth 2014挑战的一部分开发和评估的方法,以确定纵向临床叙述中由25种实体类型定义的PHI。我们的方法将知识驱动(词典和规则)和数据驱动(机器学习)方法与大量特征相结合,以解决特定命名实体的去标识问题。此外,我们设计了一种两遍识别方法,该方法根据第一步中高置信度确定的PHI实体创建特定于患者的运行时词典,然后在第二遍中使用该词典来识别缺乏特定线索的提及。该方法在测试数据集(514个叙事)上实现了91%的严格和95%的符号级评价的总体微观F1度量。虽然大多数PHI条目都可以可靠地识别,但特别具有挑战性的是提到组织和专业。尽管如此,总体结果表明,自动化文本挖掘方法可以用于可靠地处理临床记录,以识别个人信息,从而为进一步的临床和流行病学研究提供大规模去识别非结构化数据的关键一步。
A recent promise to access unstructured clinical data from electronic health records on large-scale has revitalized the interest in automated de-identification of clinical notes, which includes the identification of mentions of Protected Health Information (PHI). We describe the methods developed and evaluated as part of the i2b2/UTHealth 2014 challenge to identify PHI defined by 25 entity types in longitudinal clinical narratives. Our approach combines knowledge-driven (dictionaries and rules) and data-driven (machine learning) methods with a large range of features to address de-identification of specific named entities. In addition, we have devised a two-pass recognition approach that creates a patient-specific run-time dictionary from the PHI entities identified in the first step with high confidence, which is then used in the second pass to identify mentions that lack specific clues. The proposed method achieved the overall micro F1-measures of 91% on strict and 95% on token-level evaluation on the test dataset (514 narratives). Whilst most PHI entites can be reliably identified, particularly challenging were mentions of Organisations and Professions. Still, the overall results suggest that automated text mining methods can be used to reliably process clinical notes to identify personal information and thus providing a crucial step in large-scale de-identification of unstructured data for further clinical and epidemiological studies.