De-identification of clinical notes via recurrent neural network and conditional random field.
De-identification of clinical notes via recurrent neural network and conditional random field.
复制标题
通过循环神经网络和条件随机场对临床记录进行去识别
DOI:
10.1016/j.jbi.2017.05.023
复制
发表时间:
2017-11
影响因子:
4.5
通讯作者:
Chen Q
中科院分区:
文献类型:
--
作者:
Liu Z;Tang B;Wang X;Chen Q
De-identification, identifying information from data, such as protected health information (PHI) present in clinical data, is a critical step to enable data to be shared or published. The 2016 Centers of Excellence in Genomic Science (CEGS) Neuropsychiatric Genome-scale and RDOC Individualized Domains (N-GRID) clinical natural language processing (NLP) challenge contains a de-identification track in de-identifying electronic medical records (EMRs) (i.e., track 1). The challenge organizers provide 1000 annotated mental health records for this track, 600 out of which are used as a training set and 400 as a test set. We develop a hybrid system for the de-identification task on the training set. Firstly, four individual subsystems, that is, a subsystem based on bidirectional LSTM (long-short term memory, a variant of recurrent neural network), a subsystem-based on bidirectional LSTM with features, a subsystem based on conditional random field (CRF) and a rule-based subsystem, are used to identify PHI instances. Then, an ensemble learning-based classifiers is deployed to combine all PHI instances predicted by above three machine learning-based subsystems. Finally, the results of the ensemble learning-based classifier and the rule-based subsystem are merged together. Experiments conducted on the official test set show that our system achieves the highest micro F1-scores of 93.07%, 91.43% and 95.23% under the “token”, “strict” and “binary token” criteria respectively, ranking first in the 2016 CEGS N-GRID NLP challenge. In addition, on the dataset of 2014 i2b2 NLP challenge, our system achieves the highest micro F1-scores of 96.98%, 95.11% and 98.28% under the “token”, “strict” and “binary token” criteria respectively, outperforming other state-of-the-art systems. All these experiments prove the effectiveness of our proposed method.
登录
查看更多内容
影响因子:
4.5
作者:
He B;Guan Y;Cheng J;Cen K;Hua W
通讯作者:
Hua W
影响因子:
4.5
作者:
Dehghan A;Kovacevic A;Karystianis G;Keane JA;Nenadic G
通讯作者:
Nenadic G
影响因子:
3.5
作者:
Gupta, D;Saul, M;Gilbertson, J
通讯作者:
Gilbertson, J
影响因子:
6.8
作者:
Eddy, SR
通讯作者:
Eddy, SR
影响因子:
1.1
作者:
Freund, Y;Schapire, RE
通讯作者:
Schapire, RE