Supervised embedding of textual predictors with applications in clinical diagnostics for pediatric cardiology.

Supervised embedding of textual predictors with applications in clinical diagnostics for pediatric cardiology.
复制标题

DOI:
10.1136/amiajnl-2013-001792
复制
发表时间:
2014-02
期刊:
Journal of the American Medical Informatics Association : JAMIA
影响因子:
--
通讯作者:
Thomas Perry;H. Zha;Ke Zhou;P. Frias;Dadan Zeng;M. Braunstein
Thomas Perry;H. Zha;Ke Zhou;P. Frias;Dadan Zeng;M. Braunstein
中科院分区:
其他
文献类型:
--
作者:
Thomas Perry;H. Zha;Ke Zhou;P. Frias;Dadan Zeng;M. Braunstein

文献摘要

被引文献

相似文献

目的:电子健康记录为基于机器学习的诊断辅助工具提供关键的预测信息。然而,由于文本数据的高维性,许多传统的机器学习方法无法同时将文本数据整合到预测过程中。在本文中,我们提出了一种使用拉普拉斯特征映射的监督方法,使现有的机器学习方法能够同时估计文本数据的低维表示和基于这些低维表示的准确预测器。材料与方法提出了一种监督拉普拉斯特征映射方法,通过将文本预测器嵌入到低维潜在空间中来增强预测模型,从而保持文本数据在高维空间中的局部相似性。该算法采用梯度下降法进行交替优化。为了评估,我们将我们的方法应用于来自大型单中心儿科心脏病学实践的2000多例患者记录,以预测患者是否被诊断患有心脏病。在我们的实验中,由于数据的可用性,我们考虑相对较短的文本描述。我们将该方法与潜在语义索引、潜在Dirichlet分配和局部Fisher判别分析进行了比较。评估结果采用四项指标:受者工作特征曲线下面积(AUC)、马修斯相关系数(MCC)、特异性和敏感性。结果与讨论结果表明,监督拉普拉斯特征映射是我们研究中表现最好的方法,AUC和MCC分别达到0.782和0.374。有监督的拉普拉斯特征图显示,与排除文本数据的基线相比,AUC和MCC分别增加了8.16%和20.6%,AUC和MCC分别比无监督的拉普拉斯特征图增加了2.69%和5.35%。作为解决方案,我们提出了一种监督拉普拉斯特征映射方法,将文本预测器嵌入到低维欧几里得空间中。这种方法允许许多现有的机器学习预测器有效地捕获文本预测器的潜力,特别是那些基于短文本的预测器。
OBJECTIVE Electronic health records possess critical predictive information for machine-learning-based diagnostic aids. However, many traditional machine learning methods fail to simultaneously integrate textual data into the prediction process because of its high dimensionality. In this paper, we present a supervised method using Laplacian Eigenmaps to enable existing machine learning methods to estimate both low-dimensional representations of textual data and accurate predictors based on these low-dimensional representations at the same time. MATERIALS AND METHODS We present a supervised Laplacian Eigenmap method to enhance predictive models by embedding textual predictors into a low-dimensional latent space, which preserves the local similarities among textual data in high-dimensional space. The proposed implementation performs alternating optimization using gradient descent. For the evaluation, we applied our method to over 2000 patient records from a large single-center pediatric cardiology practice to predict if patients were diagnosed with cardiac disease. In our experiments, we consider relatively short textual descriptions because of data availability. We compared our method with latent semantic indexing, latent Dirichlet allocation, and local Fisher discriminant analysis. The results were assessed using four metrics: the area under the receiver operating characteristic curve (AUC), Matthews correlation coefficient (MCC), specificity, and sensitivity. RESULTS AND DISCUSSION The results indicate that supervised Laplacian Eigenmaps was the highest performing method in our study, achieving 0.782 and 0.374 for AUC and MCC, respectively. Supervised Laplacian Eigenmaps showed an increase of 8.16% in AUC and 20.6% in MCC over the baseline that excluded textual data and a 2.69% and 5.35% increase in AUC and MCC, respectively, over unsupervised Laplacian Eigenmaps. CONCLUSIONS As a solution, we present a supervised Laplacian Eigenmap method to embed textual predictors into a low-dimensional Euclidean space. This method allows many existing machine learning predictors to effectively and efficiently capture the potential of textual predictors, especially those based on short texts.