De-identification of clinical notes via recurrent neural network and conditional random field.

De-identification of clinical notes via recurrent neural network and conditional random field.
复制标题

通过循环神经网络和条件随机场对临床记录进行去识别

DOI:
10.1016/j.jbi.2017.05.023
复制
发表时间:
2017-11
影响因子:
4.5
通讯作者:
Chen Q
Chen Q
中科院分区:
医学3区
文献类型:
--
作者:
Liu Z;Tang B;Wang X;Chen Q

文献摘要

参考文献

被引文献

相似文献

去识别,即从数据中识别信息,例如临床数据中存在的受保护健康信息(PHI),是使数据能够共享或发布的关键步骤。2016年基因组科学卓越中心(CEGS)神经精神病学基因组规模和RDOC个性化领域(N-GRID)临床自然语言处理(NLP)挑战赛包含去识别电子病历(EMR)中的去识别轨道(即,轨道1)。挑战赛组织者为这条赛道提供了1000个带注释的心理健康记录,其中600个用作训练集,400个用作测试集。我们开发了一个混合系统的训练集上的去识别任务。首先,四个独立的子系统,即基于双向LSTM(长短期记忆,递归神经网络的变体)的子系统,基于双向LSTM的子系统,基于条件随机场(CRF)的子系统和基于规则的子系统,用于识别PHI实例。然后,部署基于集成学习的分类器来联合收割机组合由上述三个基于机器学习的子系统预测的所有PHI实例。最后,集成学习分类器和基于规则的子系统的结果合并在一起。在官方测试集上进行的实验表明,我们的系统在“token”,“strict”和“binary token”标准下分别获得了93.07%,91.43%和95.23%的最高微观F1分数,在2016年CEGS N-GRID NLP挑战赛中排名第一。此外,在2014年i2 b2 NLP挑战赛的数据集上,我们的系统在“token”,“strict”和“binary token”标准下分别获得了96.98%,95.11%和98.28%的最高微观F1分数,优于其他最先进的系统。实验结果证明了该方法的有效性。
De-identification, identifying information from data, such as protected health information (PHI) present in clinical data, is a critical step to enable data to be shared or published. The 2016 Centers of Excellence in Genomic Science (CEGS) Neuropsychiatric Genome-scale and RDOC Individualized Domains (N-GRID) clinical natural language processing (NLP) challenge contains a de-identification track in de-identifying electronic medical records (EMRs) (i.e., track 1). The challenge organizers provide 1000 annotated mental health records for this track, 600 out of which are used as a training set and 400 as a test set. We develop a hybrid system for the de-identification task on the training set. Firstly, four individual subsystems, that is, a subsystem based on bidirectional LSTM (long-short term memory, a variant of recurrent neural network), a subsystem-based on bidirectional LSTM with features, a subsystem based on conditional random field (CRF) and a rule-based subsystem, are used to identify PHI instances. Then, an ensemble learning-based classifiers is deployed to combine all PHI instances predicted by above three machine learning-based subsystems. Finally, the results of the ensemble learning-based classifier and the rule-based subsystem are merged together. Experiments conducted on the official test set show that our system achieves the highest micro F1-scores of 93.07%, 91.43% and 95.23% under the “token”, “strict” and “binary token” criteria respectively, ranking first in the 2016 CEGS N-GRID NLP challenge. In addition, on the dataset of 2014 i2b2 NLP challenge, our system achieves the highest micro F1-scores of 96.98%, 95.11% and 98.28% under the “token”, “strict” and “binary token” criteria respectively, outperforming other state-of-the-art systems. All these experiments prove the effectiveness of our proposed method.
DOI: 10.1016/j.jbi.2015.08.012
发表时间: 2015-12
影响因子: 4.5
作者:
He B;Guan Y;Cheng J;Cen K;Hua W
通讯作者: Hua W
DOI: 10.1016/j.jbi.2015.06.029
发表时间: 2015-12
影响因子: 4.5
作者:
Dehghan A;Kovacevic A;Karystianis G;Keane JA;Nenadic G
通讯作者: Nenadic G
DOI: 10.1309/e6k33gbpe5c27fyu
发表时间: 2004-02-01
影响因子: 3.5
作者:
Gupta, D;Saul, M;Gilbertson, J
通讯作者: Gilbertson, J
DOI: 10.1016/s0959-440x(96)80056-x
发表时间: 1996-06-01
影响因子: 6.8
作者:
Eddy, SR
通讯作者: Eddy, SR
DOI: 10.1006/jcss.1997.1504
发表时间: 1997-08-01
影响因子: 1.1
作者:
Freund, Y;Schapire, RE
通讯作者: Schapire, RE