Extracting laboratory test information from biomedical text.

Extracting laboratory test information from biomedical text.
复制标题

DOI:
10.4103/2153-3539.117450
复制
发表时间:
2013
影响因子:
--
通讯作者:
Kayaalp M
Kayaalp M
中科院分区:
其他
文献类型:
--
作者:
Kang YS;Kayaalp M

文献摘要

相似文献

目前的自然语言处理(NLP)方法在从叙事文件中提取实验室测试信息方面的有效性尚未有研究报道。本研究探讨了病理学信息学的问题,即使用当前的工具和技术,特别是机器学习和符号NLP方法,如何准确地从文本中提取这些信息。研究数据来自美国食品和药物管理局维护的文本语料库,其中包含丰富的实验室测试和测试设备信息。作者开发了一个符号信息提取(SIE)系统,用于提取有关四种类型实验室测试实体的设备和测试特定信息:标本,分析物,测量单位和检出限。他们比较了SIE和三个著名的基于机器学习的NLP系统(LingPipe、GATE和BANNER)的性能,每个系统分别实现了不同的监督机器学习方法、隐马尔可夫模型、支持向量机和条件随机场。机器学习系统识别的实验室测试实体具有中等高的召回率,但准确率较低。当不同实体值的数量(例如,标本的光谱)非常有限时,或者当实体的词法形态学非常独特时(如在度量单位中),它们的召回率相对较高,但SIE在提取标本、分析物和检测极限信息的精度和F-measure方面都优于它们,在统计上具有显著的边际。其高召回性能在分析物信息提取上具有统计学意义。尽管与机器学习方法相比存在缺点,但一个精心定制的符号系统可能会更好地识别一堆相同类型的信息之间的相关性,并可能通过利用词汇非本地上下文信息(如文档结构)胜过机器学习系统。
No previous study reported the efficacy of current natural language processing (NLP) methods for extracting laboratory test information from narrative documents. This study investigates the pathology informatics question of how accurately such information can be extracted from text with the current tools and techniques, especially machine learning and symbolic NLP methods. The study data came from a text corpus maintained by the U.S. Food and Drug Administration, containing a rich set of information on laboratory tests and test devices. The authors developed a symbolic information extraction (SIE) system to extract device and test specific information about four types of laboratory test entities: Specimens, analytes, units of measures and detection limits. They compared the performance of SIE and three prominent machine learning based NLP systems, LingPipe, GATE and BANNER, each implementing a distinct supervised machine learning method, hidden Markov models, support vector machines and conditional random fields, respectively. Machine learning systems recognized laboratory test entities with moderately high recall, but low precision rates. Their recall rates were relatively higher when the number of distinct entity values (e.g., the spectrum of specimens) was very limited or when lexical morphology of the entity was distinctive (as in units of measures), yet SIE outperformed them with statistically significant margins on extracting specimen, analyte and detection limit information in both precision and F-measure. Its high recall performance was statistically significant on analyte information extraction. Despite its shortcomings against machine learning methods, a well-tailored symbolic system may better discern relevancy among a pile of information of the same type and may outperform a machine learning system by tapping into lexically non-local contextual information such as the document structure.