Clinical Phrase Mining with Language Models

Clinical Phrase Mining with Language Models
复制标题

DOI:
10.1109/bibm49941.2020.9313496
复制
发表时间:
2020-12
期刊:
2020 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)
影响因子:
--
通讯作者:
Kaushik Mani;Xiang Yue;Bernal Jimenez Gutierrez;Yungui Huang;Simon M. Lin;Huan Sun
Kaushik Mani;Xiang Yue;Bernal Jimenez Gutierrez;Yungui Huang;Simon M. Lin;Huan Sun
中科院分区:
其他
文献类型:
--
作者:
Kaushik Mani;Xiang Yue;Bernal Jimenez Gutierrez;Yungui Huang;Simon M. Lin;Huan Sun

文献摘要

相似文献

大量重要的临床数据可在非结构化文本中获得,例如电子病历(EMR)中的出院总结和程序说明。将这些非结构化数据自动转换为结构化单元对于临床信息学领域的有效数据分析至关重要。识别以简洁和全面的方式揭示重要医学信息的短语是这个过程中的基本步骤。为opendomain文本构建的现有系统被设计为检测大多数非医学短语,而专门为从临床文本中提取概念而设计的工具不能扩展到大型语料库,并且经常遗漏围绕这些检测到的临床概念的基本上下文。我们通过提出一个框架CliniPhrase来解决这些问题,该框架适用于特定领域的基于深度神经网络的语言模型(如ClinicalBERT),以有效和高效地从具有有限训练数据的临床文档中提取高质量短语。MIMIC-III数据集上的实验结果表明,我们的方法在F1测量方面可以比当前最先进的技术高出18%,同时非常高效(快48倍)。11我们的源代码,预训练模型和文档可在线获得:https://github.com/kaushikmani/PhraseMiningLM
A vast amount of vital clinical data is available within unstructured texts such as discharge summaries and procedure notes in Electronic Medical Records (EMRs). Automatically transforming such unstructured data into structured units is crucial for effective data analysis in the field of clinical informatics. Recognizing phrases that reveal important medical information in a concise and thorough manner is a fundamental step in this process. Existing systems that are built for opendomain texts are designed to detect mostly non-medical phrases, while tools designed specifically for extracting concepts from clinical texts are not scalable to large corpora and often leave out essential context surrounding those detected clinical concepts. We address these issues by proposing a framework, CliniPhrase, which adapts domain-specific deep neural network based language models (such as ClinicalBERT) to effectively and efficiently extract high-quality phrases from clinical documents with a limited amount of training data. Experimental results on the MIMIC-III dataset show that our method can outperform the current state-of-the-art techniques by up to 18% in terms of F1 measure while being very efficient (up to 48 times faster).11Our source code, pre-trained models and documentations are available online at: https://github.com/kaushikmani/PhraseMiningLM