Automatic Deidentification by using Sentence Features and Label Consistency

Automatic Deidentification by using Sentence Features and Label Consistency
复制标题

DOI:
--
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
E. Aramaki;Takeshi Imai;K. Miyo;K. Ohe
E. Aramaki;Takeshi Imai;K. Miyo;K. Ohe
中科院分区:
其他
文献类型:
--
作者:
E. Aramaki;Takeshi Imai;K. Miyo;K. Ohe

文献摘要

相似文献

临床病历的非识别性问题已经引起了医学界的广泛关注。由于临床记录中的文本大多是不符合语法和支离破碎的,以前的方法只依赖于本地信息,即围绕当前目标词的上下文单词。本文提出了一种新的方法,采用三种类型的非局部特征,这不是来自周围的话:(1)句子特征,对应于前/下一个句子的信息和(2)标签一致性,倾向于相同的标签相同的单词序列。实验结果表明,高性能(准确率98.29%,召回率96.66%,F-措施97.47),证明了所提出的方法的可行性。
Deidentification of clinical records has drawn a great deal of attention in the medical field. Since texts in clinical records are mostly ungrammatical and fragmented, previous approaches have relied only on local information, namely contextual words surrounding a current target word. The present paper proposes a new approach employing three types of non-local features, which does not come from surrounding words: (1) sentence features, corresponding to the previous/next sentence information and (2) label consistency, preferring the same label for the same word sequence. The experimental results showed high performance (precision 98.29%; recall 96.66%; f-measure 97.47), demonstrating the feasibility of the proposed approach.