A survey of word embeddings for clinical text.
A survey of word embeddings for clinical text.
复制标题
DOI:
10.1016/j.yjbinx.2019.100057
复制
发表时间:
2019-01-01
影响因子:
4.5
通讯作者:
Rudzicz, Frank
中科院分区:
文献类型:
--
作者:
Khattak, Faiza Khan;Jeblee, Serena;Rudzicz, Frank
Representing words as numerical vectors based on the contexts in which they appear has become the de facto method of analyzing text with machine learning. In this paper, we provide a guide for training these representations on clinical text data, using a survey of relevant research. Specifically, we discuss different types of word representations, clinical text corpora, available pre-trained clinical word vector embeddings, intrinsic and extrinsic evaluation, applications, and limitations of these approaches. This work can be used as a blueprint for clinicians and healthcare workers who may want to incorporate clinical text features in their own models and applications.