Completing Missing Prevalence Rates for Multiple Chronic Diseases by Jointly Leveraging Both Intra- and Inter-Disease Population Health Data Correlations

Completing Missing Prevalence Rates for Multiple Chronic Diseases by Jointly Leveraging Both Intra- and Inter-Disease Population Health Data Correlations
复制标题

DOI:
10.1145/3442381.3449811
复制
发表时间:
2021-04
期刊:
Proceedings of the Web Conference 2021
影响因子:
--
通讯作者:
Yujie Feng;Jiangtao Wang;Yasha Wang;A. Helal
Yujie Feng;Jiangtao Wang;Yasha Wang;A. Helal
中科院分区:
其他
文献类型:
--
作者:
Yujie Feng;Jiangtao Wang;Yasha Wang;A. Helal

文献摘要

被引文献

相似文献

人口健康数据在互联网上比以往任何时候都更加公开。这些数据集为更好地了解人口的健康状况提供了巨大的潜力,并为卫生专业人员和决策者提供信息,以更好地规划资源,疾病管理和预防不同区域。然而,由于收集这些公共卫生数据的费力和高成本性质,在这些数据集上发现许多缺失条目是常见的,这对数据的实用性提出了挑战,并阻碍了可靠的分析和理解。为了解决这个问题,本文提出了一种基于深度学习的方法,称为压缩人口健康(CPH),以推断和恢复(以完成)多种慢性病的缺失患病率条目。CPH的关键见解依赖于疾病内和疾病间相关性机会的综合利用。具体来说,我们首先提出了一种基于卷积神经网络(CNN)的方法来提取和建模这两种类型的相关性,然后采用基于生成对抗网络(GAN)的流行率推理模型来联合融合它们,以便于缺失条目的流行率数据恢复。我们广泛评估的推理模型的基础上公开在Web上的真实世界的公共卫生数据集。结果表明,我们的推理方法优于其他基线方法在各种设置和显着提高的准确性(从14.8%到9.1%)。
Population health data are becoming more and more publicly available on the Internet than ever before. Such datasets offer a great potential for enabling a better understanding of the health of populations, and inform health professionals and policy makers for better resource planning, disease management and prevention across different regions. However, due to the laborious and high-cost nature of collecting such public health data, it is a common place to find many missing entries on these datasets, which challenges the utility of the data and hinders reliable analysis and understanding. To tackle this problem, this paper proposes a deep-learning-based approach, called Compressive Population Health (CPH), to infer and recover (to complete) the missing prevalence rate entries of multiple chronic diseases. The key insight of CPH relies on the combined exploitation of both intra-disease and inter-disease correlation opportunities. Specifically, we first propose a Convolutional Neural Network (CNN) based approach to extract and model both of these two types of correlations, and then adopt a Generative Adversarial Network (GAN) based prevalence inference model to jointly fuse them to facility the prevalence rates data recovery of missing entries. We extensively evaluate the inference model based on real-world public health datasets publicly available on the Web. Results show that our inference method outperforms other baseline methods in various settings and with a significantly improved accuracy (from 14.8% to 9.1%).