Recognition and Evaluation of Clinical Section Headings in Clinical Documents Using Token-Based Formulation with Conditional Random Fields.

Recognition and Evaluation of Clinical Section Headings in Clinical Documents Using Token-Based Formulation with Conditional Random Fields.
复制标题

DOI:
10.1155/2015/873012
复制
发表时间:
2015
影响因子:
--
通讯作者:
Wu CC
Wu CC
中科院分区:
生物学3区
文献类型:
--
作者:
Dai HJ;Syed-Abdul S;Chen CW;Wu CC

文献摘要

被引文献

相似文献

电子健康记录(EHR)是一种数字数据格式,收集有关单个患者或人群的电子健康信息。为了提高EHR的有意义的使用,信息提取技术已被开发出来,以识别EHR中提到的临床概念。然而,EHR的临床判断不能仅仅基于所识别的概念而不考虑其上下文信息。为了提高电子病历的可读性和可访问性,本工作开发了一个用于临床文档的章节标题识别系统。与将章节标题识别任务制定为句子分类问题相比,这项工作提出了一种基于条件随机场(CRF)模型的标记公式。一个标准的章节标题识别语料库,由具有临床经验的注释者编写,以评估性能,并将其与句子分类和基于词典的方法进行比较。实验结果表明,该方法取得了令人满意的F值为0.942,优于基于词典的方法和最好的基于词典的系统分别为0.087和0.096。我们的公式相对于基于标题的方法的一个重要优势是,它提供了一个集成的解决方案,而不需要开发额外的语法规则来将标题与周围的章节内容隔离开来。
Electronic health record (EHR) is a digital data format that collects electronic health information about an individual patient or population. To enhance the meaningful use of EHRs, information extraction techniques have been developed to recognize clinical concepts mentioned in EHRs. Nevertheless, the clinical judgment of an EHR cannot be known solely based on the recognized concepts without considering its contextual information. In order to improve the readability and accessibility of EHRs, this work developed a section heading recognition system for clinical documents. In contrast to formulating the section heading recognition task as a sentence classification problem, this work proposed a token-based formulation with the conditional random field (CRF) model. A standard section heading recognition corpus was compiled by annotators with clinical experience to evaluate the performance and compare it with sentence classification and dictionary-based approaches. The results of the experiments showed that the proposed method achieved a satisfactory F-score of 0.942, which outperformed the sentence-based approach and the best dictionary-based system by 0.087 and 0.096, respectively. One important advantage of our formulation over the sentence-based approach is that it presented an integrated solution without the need to develop additional heuristics rules for isolating the headings from the surrounding section contents.