A survey of word embeddings for clinical text.

A survey of word embeddings for clinical text.
复制标题

DOI:
10.1016/j.yjbinx.2019.100057
复制
发表时间:
2019-01-01
影响因子:
4.5
通讯作者:
Rudzicz, Frank
Rudzicz, Frank
中科院分区:
医学3区
文献类型:
--
作者:
Khattak, Faiza Khan;Jeblee, Serena;Rudzicz, Frank

文献摘要

被引文献

相似文献

根据单词出现的上下文将单词表示为数值向量已成为使用机器学习分析文本的事实上的方法。在本文中,我们通过相关研究的调查,提供了在临床文本数据上训练这些表示的指南。具体来说,我们讨论了不同类型的单词表示、临床文本语料库、可用的预训练临床词向量嵌入、内在和外在评估、应用以及这些方法的局限性。这项工作可以作为临床医生和医疗保健工作者的蓝图,他们可能希望将临床文本特征纳入自己的模型和应用程序中。
Representing words as numerical vectors based on the contexts in which they appear has become the de facto method of analyzing text with machine learning. In this paper, we provide a guide for training these representations on clinical text data, using a survey of relevant research. Specifically, we discuss different types of word representations, clinical text corpora, available pre-trained clinical word vector embeddings, intrinsic and extrinsic evaluation, applications, and limitations of these approaches. This work can be used as a blueprint for clinicians and healthcare workers who may want to incorporate clinical text features in their own models and applications.