Speculation detection for Chinese clinical notes: Impacts of word segmentation and embedding models.

Speculation detection for Chinese clinical notes: Impacts of word segmentation and embedding models.
复制标题

中文临床笔记的推测检测:分词和嵌入模型的影响

DOI:
10.1016/j.jbi.2016.02.011
复制
发表时间:
2016-04
影响因子:
4.5
通讯作者:
Lei J
Lei J
中科院分区:
医学3区
文献类型:
--
作者:
Zhang S;Kang T;Zhang X;Wen D;Elhadad N;Lei J

文献摘要

相似文献

推测代表了对某些事实的不确定性。在临床文本中,识别推测是自然语言处理(NLP)的关键步骤。虽然在许多语言中这是一项不平凡的任务,但检测中文临床笔记中的推测可能特别具有挑战性,因为分词可能是必要的上游操作。本文的目标是构建一个最先进的中文临床笔记推测检测系统,并研究嵌入特征和分词是否值得利用来完成这一总体任务。我们提出了一个基于序列标记的投机检测系统,它依赖于字符袋,字袋,字符嵌入和字嵌入的功能。我们在一个新的数据集上进行实验,该数据集包含36,828个临床笔记,其中2,000个笔记上有5,103个黄金标准推测注释,并比较了基于一般和特定领域分割器分别给出的单词分割计算单词嵌入的系统。我们的系统能够达到高达92.2%的F分数测量的性能。我们证明,分词是至关重要的,以产生高质量的词嵌入,以促进下游的信息提取应用程序,并建议,一个域相关的分词器可以是至关重要的,这样的临床自然语言处理任务在中文。
Speculations represent uncertainty towards certain facts. In clinical texts, identifying speculations is a critical step of natural language processing (NLP). While it is a nontrivial task in many languages, detecting speculations in Chinese clinical notes can be particularly challenging because word segmentation may be necessary as an upstream operation. The objective of this paper is to construct a state-of-the-art speculation detection system for Chinese clinical notes and to investigate whether embedding features and word segmentations are worth exploiting towards this overall task. We propose a sequence labeling based system for speculation detection, which relies on features from bag of characters, bag of words, character embedding, and word embedding. We experiment on a novel dataset of 36,828 clinical notes with 5,103 gold-standard speculation annotations on 2,000 notes, and compare the systems in which word embeddings are calculated based on word segmentations given by general and by domain specific segmenters respectively. Our systems are able to reach performance as high as 92.2% measured by F score. We demonstrate that word segmentation is critical to produce high quality word embedding to facilitate downstream information extraction applications, and suggest that a domain dependent word segmenter can be vital to such a clinical NLP task in Chinese language.