Using phrases and document metadata to improve topic modeling of clinical reports.

Using phrases and document metadata to improve topic modeling of clinical reports.
复制标题

DOI:
10.1016/j.jbi.2016.04.005
复制
发表时间:
2016-06
影响因子:
4.5
通讯作者:
Arnold CW
Arnold CW
中科院分区:
医学3区
文献类型:
--
作者:
Speier W;Ong MK;Arnold CW

文献摘要

相似文献

概率主题模型为分析非结构化文本提供了一种无监督的方法,具有集成到临床自动摘要系统中的潜力。临床文档伴随着患者病史中的元数据,并且经常包含多个词的概念,这些概念对于准确地解释所包含的文本是有价值的。虽然现有的方法试图单独解决这些问题,但我们为自由文本临床文档提供了一个统一的模型,该模型集成了上下文患者和文档级数据,并发现了多个单词的概念。在该模型中,短语由链式n元语法表示,Dirichlet超参数根据文档级和患者级上下文进行加权。该方法和其他三个潜在的Dirichlet分配模型适用于大量的临床报告集合。结果主题的例子展示了新模型的结果,并使用经验对数似然对表示的质量进行了评估。提出的模型能够基于患者和文档信息创建信息性先验概率,并捕获代表各种临床概念的短语。与比较的方法相比,使用该模型的表示具有显著更高的经验对数似然。整合文档元数据和捕获临床文本中的短语大大提高了临床文档的主题表示。由此产生的临床信息主题可以有效地作为临床报告自动摘要系统的基础。
Probabilistic topic models provide an unsupervised method for analyzing unstructured text, which have the potential to be integrated into clinical automatic summarization systems. Clinical documents are accompanied by metadata in a patient’s medical history and frequently contains multiword concepts that can be valuable for accurately interpreting the included text. While existing methods have attempted to address these problems individually, we present a unified model for free-text clinical documents that integrates contextual patient- and document-level data, and discovers multi-word concepts. In the proposed model, phrases are represented by chained n-grams and a Dirichlet hyper-parameter is weighted by both document-level and patient-level context. This method and three other Latent Dirichlet allocation models were fit to a large collection of clinical reports. Examples of resulting topics demonstrate the results of the new model and the quality of the representations are evaluated using empirical log likelihood. The proposed model was able to create informative prior probabilities based on patient and document information, and captured phrases that represented various clinical concepts. The representation using the proposed model had a significantly higher empirical log likelihood than the compared methods. Integrating document metadata and capturing phrases in clinical text greatly improves the topic representation of clinical documents. The resulting clinically informative topics may effectively serve as the basis for an automatic summarization system for clinical reports.