Multi-task Learning for Aggregated Data using Gaussian Processes

Multi-task Learning for Aggregated Data using Gaussian Processes
复制标题

DOI:
--
复制
发表时间:
2019-06
期刊:
ArXiv
影响因子:
--
通讯作者:
F. Yousefi;M. Smith;Mauricio A Álvarez
F. Yousefi;M. Smith;Mauricio A Álvarez
中科院分区:
其他
文献类型:
--
作者:
F. Yousefi;M. Smith;Mauricio A Álvarez

文献摘要

被引文献

相似文献

汇总数据在流行病学和人口学等领域很常见。例如,人口普查数据通常以时间段或空间分辨率(城市、地区或国家)定义的平均值给出。在本文中,我们提出了一种基于高斯过程的新颖的多任务学习模型,用于对在不同输入尺度上聚合的变量进行联合学习。我们的模型将每个任务表示为每个任务以不同规模集成的潜在流程实现的线性组合。然后,我们能够通过分析或数值方式计算不同任务之间的互协方差。我们还允许每个任务具有潜在不同的似然模型,并提供可以以随机方式优化的变分下界,使我们的模型适用于更大的数据集。我们在综合示例、生育率数据集和空气污染预测应用程序中展示了该模型的示例。
Aggregated data is commonplace in areas such as epidemiology and demography. For example, census data for a population is usually given as averages defined over time periods or spatial resolutions (cities, regions or countries). In this paper, we present a novel multi-task learning model based on Gaussian processes for joint learning of variables that have been aggregated at different input scales. Our model represents each task as the linear combination of the realizations of latent processes that are integrated at a different scale per task. We are then able to compute the cross-covariance between the different tasks either analytically or numerically. We also allow each task to have a potentially different likelihood model and provide a variational lower bound that can be optimised in a stochastic fashion making our model suitable for larger datasets. We show examples of the model in a synthetic example, a fertility dataset, and an air pollution prediction application.