Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus.

Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus.
复制标题

DOI:
10.1016/j.jbi.2015.07.020
复制
发表时间:
2015-12
影响因子:
4.5
通讯作者:
Uzuner Ö
Uzuner Ö
中科院分区:
医学3区
文献类型:
--
作者:
Stubbs A;Uzuner Ö

文献摘要

被引文献

相似文献

2014年i2 b2/UTHealth自然语言处理共享任务的重点是纵向医疗记录的去识别化。在这条追踪中,我们对描述296名患者的1,304份纵向医疗记录进行了去识别。该语料库在对HIPAA指南的广义解释下使用双注释进行去标识,然后进行仲裁、多轮健全性检查和证明阅读。与黄金标准相比,注释器基于令牌的平均F1度量为0.927。由此产生的注释用于对数据进行去识别,并为2014年i2 b2/UTHealth共享任务的去识别跟踪设定了黄金标准。所有带注释的私人健康信息都被自动替换为现实的替代品,然后人工阅读和纠正。由此产生的语料库是第一个可用于去识别研究的同类语料库。该语料库首次用于2014年i2 b2/UTHealth共享任务,在此期间,系统使用基于实体的微平均评估实现了平均F测量值0.872和最大F测量值0.964。
The 2014 i2b2/UTHealth natural language processing shared task featured a track focused on the de-identification of longitudinal medical records. For this track, we de-identified a set of 1,304 longitudinal medical records describing 296 patients. This corpus was de-identified under a broad interpretation of the HIPAA guidelines using double-annotation followed by arbitration, rounds of sanity checking, and proof reading. The average token-based F1 measure for the annotators compared to the gold standard was 0.927. The resulting annotations were used both to de-identify the data and to set the gold standard for the de-identification track of the 2014 i2b2/UTHealth shared task. All annotated private health information were replaced with realistic surrogates automatically and then read over and corrected manually. The resulting corpus is the first of its kind made available for de-identification research. This corpus was first used for the 2014 i2b2/UTHealth shared task, during which the systems achieved a mean F-measure of 0.872 and a maximum F-measure of 0.964 using entity-based micro-averaged evaluations.