Representing information in patient reports using natural language processing and the extensible markup language

Representing information in patient reports using natural language processing and the extensible markup language
复制标题

DOI:
10.1136/jamia.1999.0060076
复制
发表时间:
1999-01-01
影响因子:
6.4
通讯作者:
Liu, HF
Liu, HF
中科院分区:
管理学2区
文献类型:
--
作者:
Friedman, C;Hripcsak, G;Liu, HF

文献摘要

被引文献

相似文献

目的:设计一种文档模型,为广泛的临床应用提供对患者报告中临床信息的可靠、高效的访问,并实现一种使用自然语言处理的自动化方法,将文本报告映射到与模型一致的形式。方法:使用可扩展标记语言(XML)设计了一种文档模型,该模型在保留原始内容的同时编码患者报告中的结构化临床信息,并创建文档类型定义(DTD)。对现有的自然语言处理器(NLP)进行了修改,以生成与该模型一致的输出。使用改进的NLP系统处理200份报告,并使用XML验证解析器对生成的XML输出进行验证。结果:改进的NLP系统成功处理了全部200份报告。一个报告的输出无效,199个报告是与DTD一致的有效XML格式。结论:自然语言处理可用于自动创建包含结构化组件的丰富文档,该组件的元素链接到原始文本报告的部分。这种集成的文档模型提供了一种表示,其中包含特定信息的文档可以通过查询结构化组件来准确和高效地检索。如果需要手动审查文件,还可以确定并突出显示原始报告中的重要信息。使用标记的XML模型提供了一个额外的好处,那就是操作XML文档的软件工具随时可用。
Objective: To design a document model that provides reliable and efficient access to clinical information in patient reports for a broad range of clinical applications, and to implement an automated method using natural language processing that maps textual reports to a form consistent with the model.Methods: A document model that encodes structured clinical information in patient reports while retaining the original contents was designed using the extensible markup Language (XML), and a document type definition (DTD) was created. An existing natural language processor (NLP)was modified to generate output consistent with the model. Two hundred reports were processed using the modified NLP system, and the XML output that was generated was validated using an XML validating parser.Results: The modified NLP system successfully processed all 200 reports. The output of one report was invalid, and 199 reports were valid XML forms consistent with the DTD.Conclusions: Natural language processing can be used to automatically create an enriched document that contains a structured component whose elements are linked to portions of the original textual report. This integrated document model provides a representation where documents containing specific information can be accurately and efficiently retrieved by querying the structured components. If manual review of the documents is desired, the salient information in the original reports can also be identified and highlighted. Using an XML model of tagging provides an additional benefit in that software tools that manipulate XML documents are readily available.