Deep Multimodality Model for Multi-task Multi-view Learning

Deep Multimodality Model for Multi-task Multi-view Learning
复制标题

DOI:
10.1137/1.9781611975673.2
复制
发表时间:
2019-01
期刊:
--
影响因子:
--
通讯作者:
Lecheng Zheng;Yu Cheng;Jingrui He
Lecheng Zheng;Yu Cheng;Jingrui He
中科院分区:
其他
文献类型:
--
作者:
Lecheng Zheng;Yu Cheng;Jingrui He

文献摘要

被引文献

相似文献

许多现实世界的问题表现出多种类型的异质性的共存,例如视图异质性(即,多视图属性)和任务异质性(即,多任务特性)。例如,在包含同一对象的多个姿态的图像分类问题中,每个姿态可以被认为是一个视图,并且每种类型的对象的检测可以被视为一个任务。此外,在某些问题中,多个视图的数据类型可能不同。例如,在Web分类问题中,我们可能会提供一个图像和文本混合数据集,其中网页的特征是图像和文本。解决这类问题的一个常见策略是利用视图的一致性和任务的相关性来构建预测模型。在深度神经网络的背景下,多任务相关性通常通过在每一层对任务进行分组来实现,而多视图一致性通常通过找到视图之间的最大相关系数来实现。然而,还没有现有的深度学习算法联合建模任务和视图的双重异质性,特别是对于具有多种模态的数据集(文本和图像混合数据集或文本和视频混合数据集等)。在本文中,我们通过提出一个深度多任务多视图学习框架来弥合这一差距,该框架可以学习这种双重异质性问题的深度表示。在多个真实数据集上的实验研究证明了我们提出的Deep-MTMV算法的有效性。
Many real-world problems exhibit the coexistence of multiple types of heterogeneity, such as view heterogeneity (i.e., multi-view property) and task heterogeneity (i.e., multi-task property). For example, in an image classification problem containing multiple poses of the same object, each pose can be considered as one view, and the detection of each type of object can be treated as one task. Furthermore, in some problems, the data type of multiple views might be different. In a web classification problem, for instance, we might be provided an image and text mixed data set, where the web pages are characterized by both images and texts. A common strategy to solve this kind of problem is to leverage the consistency of views and the relatedness of tasks to build the prediction model. In the context of deep neural network, multi-task relatedness is usually realized by grouping tasks at each layer, while multi-view consistency is usually enforced by finding the maximal correlation coefficient between views. However, there is no existing deep learning algorithm that jointly models task and view dual heterogeneity, particularly for a data set with multiple modalities (text and image mixed data set or text and video mixed data set, etc.). In this paper, we bridge this gap by proposing a deep multi-task multi-view learning framework that learns a deep representation for such dual-heterogeneity problems. Empirical studies on multiple real-world data sets demonstrate the effectiveness of our proposed Deep-MTMV algorithm.