Use of general-purpose negation detection to augment concept indexing of medical documents: A quantitative study using the UMLS

Use of general-purpose negation detection to augment concept indexing of medical documents: A quantitative study using the UMLS
复制标题

DOI:
10.1136/jamia.2001.0080598
复制
发表时间:
2001-11-01
影响因子:
6.4
通讯作者:
Nadkarni, PM
Nadkarni, PM
中科院分区:
管理学2区
文献类型:
--
作者:
Mutalik, PG;Deshpande, A;Nadkarni, PM

文献摘要

被引文献

相似文献

目的:为了验证这一假设,即口述医学文档中的大多数否定概念的实例可以通过依赖于用于解析形式化(计算机)语言-具体地,使用正则表达式来生成有限状态机的词法扫描器(“lexer”),以及依赖于上下文无关语法的受限子集(称为LALR(1)语法)的解析器。一个不同的训练集的40个医疗文件,从各种专业进行手动检查,并用于开发一个程序(Negfinder),其中包含的规则,以识别一个大的否定模式发生在文本中。Negfinder的词法分析器和解析器是使用通常用于生成编程语言编译器的工具开发的。Negfinder的输入包括经过预处理以识别UMLS概念的医学叙述:已识别概念的文本已被替换为包括其UMLS概念ID的编码表示。该程序生成一个索引,其中每个文档中的概念实例有一个条目,记录该概念的否定是否存在。这些信息被用来通过颜色编码标记每个文档的文本,以便更容易检查。然后用两种方法对解析器进行评估:1)60个文档的测试集(30份出院小结,30份手术记录)进行目视检查以量化假阳性和假阴性结果; 2)由人类观察者和Negfinder独立检查1个不同的测试集的10份文件的阴性,并比较结果。在使用标记文档的第一次评估中,在60个文档中检测到8,358个UMLS概念实例,其中544个是由程序检测到的否定,并通过人类观察进行验证(真阳性结果,或TP)。13个实例被错误地标记为否定(假阳性结果或FP),该程序错过了27个否定实例(假阴性结果或FN),灵敏度为95.3%,特异性为97.7%。在使用独立否定检测的第二次评估中,在10个文档中检测到1,869个概念,其中135个TP,12个Fps和6个FN,灵敏度为95.7%,特异性为91.8%。其中一个字“不”,“否认/否认”,“不”,或“没有”是目前在92.5%的所有negations.Conclusions:否定的大多数概念,在医学叙事可以可靠地检测到一个简单的策略。检测的可靠性取决于几个因素,最重要的是概念匹配的准确性。
Objectives: To test the hypothesis that most instances of negated concepts in dictated medical documents can be detected by a strategy that relies on tools developed for the parsing of formal (computer) languages-specifically, a lexical scanner ("lexer") that uses regular expressions to generate a finite state machine, and a parser that relies on a restricted subset of context-free grammars, known as LALR(1) grammars.Methods: A diverse training set of 40 medical documents from a variety of specialties was manually inspected and used to develop a program (Negfinder) that contained rules to recognize a large set of negated patterns occurring in the text. Negfinder's lexer and parser were developed using tools normally used to generate programming language compilers. The input to Negfinder consisted of medical narrative that was preprocessed to recognize UMLS concepts: the text of a recognized concept had been replaced with a coded representation that included its UMLS concept ID. The program generated an index with one entry per instance of a concept in the document, where the presence or absence of negation of that concept was recorded. This information was used to mark up the text of each document by color-coding it to make it easier to inspect. The parser was then evaluated in two ways: 1) a test set of 60 documents (30 discharge summaries, 30 surgical notes) marked-up by Negfinder was inspected visually to quantify false-positive and false-negative results; and 2) 1 different test set of 10 documents was independently examined for negatives by a human observer and by Negfinder, and the results were compared.Results: In the first evaluation using marked-up documents, 8,358 instances of UMLS concepts were detected in the 60 documents, of which 544 were negations detected by the program and verified by human observation (true-positive results, or TPs). Thirteen instances were wrongly flagged as negated (false-positive results, or FPs), and the program missed 27 instances of negation (false-negative results, or FNs), yielding a sensitivity of 95.3 percent and a specificity of 97.7 percent. In the second evaluation using independent negation detection, 1,869 concepts were detected in 10 documents, with 135 TPs, 12 Fps, and 6 FNs, yielding a sensitivity of 95.7 percent and a specificity of 91.8 percent. One of the words "no," "denies/denied," "not," or "without" was present in 92.5 percent of all negations.Conclusions: Negation of most concepts in medical narrative can be reliably detected by a simple strategy. The reliability of detection depends on several factors, the most important being the accuracy of concept matching.