Regularizing Topic Discovery in EMRs with Side Information by Using Hierarchical Bayesian Models

Regularizing Topic Discovery in EMRs with Side Information by Using Hierarchical Bayesian Models
复制标题

使用分层贝叶斯模型通过辅助信息规范 EMR 中的主题发现

DOI:
10.1109/icpr.2014.234
复制
发表时间:
2014
期刊:
2014 22nd International Conference on Pattern Recognition
影响因子:
--
通讯作者:
S. Venkatesh
S. Venkatesh
中科院分区:
--
文献类型:
--
作者:
Cheng Li;Santu Rana;Dinh Q. Phung;S. Venkatesh

文献摘要

被引文献

相似文献

我们提出了一种新颖的分层贝叶斯框架,即依赖于单词距离的中餐馆特许经营权(wd-dCRF),用于从通过单词与单词关系形式的辅助信息规范化的文档语料库中发现主题,并将其应用于电子病历(EMR)。通常,EMR 数据集由多个患者(文档)组成,每个患者包含许多诊断代码(单词)。我们利用诊断代码中以语义树结构形式提供的辅助信息来进行语义一致的疾病主题发现。当辅助信息以树结构的形式提供时,我们引入了新的函数来计算单词到单词的距离。我们使用 MCMC 技术推导出 wddCRF 的有效推理方法。我们对包含约 1000 名多血管疾病患者的真实世界医疗数据集进行评估。与流行的主题分析工具分层狄利克雷过程(HDP)相比,我们的模型发现的主题在定性和定量方面都具有优势。
We propose a novel hierarchical Bayesian framework, word-distance-dependent Chinese restaurant franchise (wd-dCRF) for topic discovery from a document corpus regularized by side information in the form of word-to-word relations, with an application on Electronic Medical Records (EMRs). Typically, a EMRs dataset consists of several patients (documents) and each patient contains many diagnosis codes (words). We exploit the side information available in the form of a semantic tree structure among the diagnosis codes for semantically-coherent disease topic discovery. We introduce novel functions to compute word-to-word distances when side information is available in the form of tree structures. We derive an efficient inference method for the wddCRF using MCMC technique. We evaluate on a real world medical dataset consisting of about 1000 patients with PolyVascular disease. Compared with the popular topic analysis tool, hierarchical Dirichlet process (HDP), our model discovers topics which are superior in terms of both qualitative and quantitative measures.