Evaluation of a Method to Identify and Categorize Section Headers in Clinical Documents

Evaluation of a Method to Identify and Categorize Section Headers in Clinical Documents
复制标题

DOI:
10.1197/jamia.m3037
复制
发表时间:
2009-11-01
影响因子:
6.4
通讯作者:
Miller, Randolph A.
Miller, Randolph A.
中科院分区:
管理学2区
文献类型:
--
作者:
Denny, Joshua C.;Spickard, Anderson, III;Miller, Randolph A.

文献摘要

被引文献

相似文献

目的:临床记录通常用自然语言书写,通常包含将其划分为部分的子结构,例如“当前疾病史”或“家庭病史”。作者设计并评估了一种算法("SecTag"),用于识别"历史和体检"文档("H & P笔记")中的标记和未标记(隐含)笔记部分标题。设计:SecTag算法使用自然语言处理技术、具有拼写校正的单词变体识别、基于术语的规则和朴素贝叶斯评分方法的组合来识别笔记部分标题。11位医生评估了SecTag在319个随机选择的H & P notes.Measures上的性能:主要结果是算法在识别所有文档部分和预定义的29个主要部分列表方面的召回率和精度。一个次要的结果是评估该算法的能力,以识别正确的开始和结束boundaries of identified sections.Results的边界:SecTag算法确定了16,036个总部分和7,858个主要部分。医生评价者将1.5,329个切片分类为真阳性,并确定了SecTag遗漏的160个切片。SecTag算法的召回率和准确率分别为99.0%和95.6%的所有部分,98.6%和96.2%的主要部分,96.6%和86.8%的未标记的部分。该算法确定了正确的开始和结束文本边界为94.8%的标记部分和85.9%的unlabeled sections.Conclusions:SecTag算法准确地确定了标记和未标记的部分在历史和物理文档。这种类型的算法可以帮助自然语言处理应用,例如临床决策支持系统或医学学员的能力评估。美国医学信息协会杂志2009; 16:806 - 815。DOI 10.1197/jamia.M3037.
Objective: Clinical notes, typically written in natural language, often contain substructure that divides them into sections, such as "History of Present Illness" or "Family Medical History." The authors designed and evaluated an algorithm ("SecTag") to identify both labeled and unlabeled (implied) note section headers in "history and physical examination" documents ("H&P notes").Design: The SecTag algorithm uses a combination of natural language processing techniques, word variant recognition with spelling correction, terminology-based rules, and naive Bayesian scoring methods to identify note section headers. Eleven physicians evaluated SecTag's performance on 319 randomly chosen H&P notes.Measurements: The primary outcomes were the algorithm's recall and precision in identifying all document sections and a predefined list of twenty-nine major sections. A secondary outcome was to evaluate the algorithm's ability to recognize the correct start and end boundaries of identified sections.Results: The SecTag algorithm identified 16,036 total sections and 7,858 major sections. Physician evaluators classified 1.5,329 as true positives and identified 160 sections omitted by SecTag. The recall and precision of the SecTag algorithm were 99.0 and 95.6% for all sections, 98.6 and 96.2% for major sections, and 96.6 and 86.8% for unlabeled sections. The algorithm determined the correct starting and ending text boundaries for 94.8% of labeled sections and 85.9% of unlabeled sections.Conclusions: The SecTag algorithm accurately identified both labeled and unlabeled sections in history and physical documents. This type of algorithm may assist in natural language processing applications, such as clinical decision support systems or competency assessment for medical trainees. J Am Med Inform Assoc. 2009;16:806-815. DOI 10.1197/jamia.M3037.