Automatic de-identification of electronic medical records using token-level and character-level conditional random fields.

Automatic de-identification of electronic medical records using token-level and character-level conditional random fields.
复制标题

DOI:
10.1016/j.jbi.2015.06.009
复制
发表时间:
2015-12
影响因子:
4.5
通讯作者:
Zhu S
Zhu S
中科院分区:
医学3区
文献类型:
--
作者:
Liu Z;Chen Y;Tang B;Wang X;Chen Q;Li H;Wang J;Deng Q;Zhu S

文献摘要

被引文献

相似文献

去识别,识别和删除包括电子病历(EMR)在内的临床数据中存在的所有受保护的健康信息(PHI),是公开临床数据的关键步骤。2014年i2 b2(整合生物学和床边信息学中心)临床自然语言处理(NLP)挑战赛为去识别化(第1道)设置了一个赛道。在这项研究中,我们提出了一个混合系统的基础上,机器学习和规则的方法去识别的轨道。在我们的系统中,PHI实例首先由两个(令牌级和字符级)条件随机场(CRF)和基于规则的分类器识别,然后通过一些规则合并。在i2 b2语料库上进行的实验表明,我们的系统在“token”,“strict”和“relaxed”标准下分别获得了94.64%,91.24%和91.63%的最高微F分数,这是2014年i2 b2挑战赛中排名最高的系统之一。在集成了一些改进的本地化词典后,我们的系统得到了进一步的改进,在“令牌”,“严格”和“宽松”标准下的F-分数分别为94.83%,91.57%和91.95%。
De-identification, identifying and removing all protected health information (PHI) present in clinical data including electronic medical records (EMRs), is a critical step in making clinical data publicly available. The 2014 i2b2 (Center of Informatics for Integrating Biology and Bedside) clinical natural language processing (NLP) challenge sets up a track for de-identification (track 1). In this study, we propose a hybrid system based on both machine learning and rule approaches for the de-identification track. In our system, PHI instances are first identified by two (token-level and character-level) conditional random fields (CRFs) and a rule-based classifier, and then are merged by some rules. Experiments conducted on the i2b2 corpus show that our system submitted for the challenge achieves the highest micro F-scores of 94.64%, 91.24% and 91.63% under the “token”, “strict” and “relaxed” criteria respectively, which is among top-ranked systems of the 2014 i2b2 challenge. After integrating some refined localization dictionaries, our system is further improved with F-scores of 94.83%, 91.57% and 91.95% under the “token”, “strict” and “relaxed” criteria respectively.