课题基金 / 基金详情

CAREER: New Directions in Deep Representation Learning from Complex Multimodal Data

CAREER: New Directions in Deep Representation Learning from Complex Multimodal Data
职业:复杂多模态数据深度表示学习的新方向
批准号:
1453651
负责人:
Honglak Lee
金额:
$48.86万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2023-08-31

项目摘要

项目成果

Honglak Lee的其他基金

相似基金

相关文献

中文摘要
翻译
深度学习的目标是学习具有层次和组合结构的数据的抽象表示。深度学习方法可以有效地从高维输入数据(例如,用于分类),并已成功应用于许多现实世界的问题,如图像分类,语音识别和文本建模。尽管取得了这些成功,但仍然存在一个具有挑战性的开放问题:我们如何才能学习一个强大的深度表示,以便从复杂数据中进行整体理解和高级推理? 这个CAREER项目旨在解决这个问题,并有望为深度表示学习中的推理、学习和优化带来新的深度架构、图形模型和算法进步。 研究成果将通过出版物、讲座和教程传播。除了推进深度学习及其所需的许多应用领域的最新发展外,该项目还将通过以下方式整合研究和教育:1)开发机器学习课程,其中将深度学习作为一个关键主题; 2)指导重要的研究生和本科生研究活动; 3)通过举办演示会议和指导科学展览/研究项目来接触K-12学生。 该项目研究了以下密切相关和互补的重点:首先,它开发了深度学习算法,以从复杂数据中分离出变化因素。这是通过用深度生成模型(例如,通过与身份、观点和情感相对应的潜在因素的交互来对面部图像进行建模)。除了更好的概括,这种方法也适用于高级推理,例如进行类比。高阶交互的建模将通过学习每个变异因子的子流形来实现,其中对应信息用于正则化潜在表示。该项目还将开发弱监督和半监督的解开算法,自动建立对应关系,而无需人工监督。其次,该项目开发了用于结构化预测问题的深度表示学习方法。具体来说,它将开发一个具有深度表示的图形模型,可以对输出变量之间的复杂依赖关系进行建模。这个框架也可以被看作是数据驱动的建模高阶先验结构化数据,并可用于建模高阶条件随机场,允许有效的推理和学习。此外,该项目还为涉及不确定性的结构化预测问题(即,一对多映射)。 第三,该项目开发了新的深度学习算法,用于从多种异构输入模式(如图像和文本、音频和视频以及多个传感器流)中构建共享表示。其主要思想是单独建模的条件分布的每个输入模态给定其他模态。这种方法解决了众所周知的困难建模跨异质多模态输入的联合分布,并提供了一个理论分析的条件下,该方法可以恢复一致的生成模型。这种提法允许强大的识别和高层次的推理,从异构多模态数据。总的来说,这三个方面是互补的,预计将在解决更广泛的人工智能问题和超越当前最先进的深度学习方面发挥协同作用。
英文摘要
The goal of deep learning is to learn an abstract representation of data with a hierarchical and compositional structure. Deep learning methods can effectively learn discriminative features from high-dimensional input data (e.g., for classification), and have been successfully applied to many real-world problems, such as image classification, speech recognition, and text modeling. Despite these successes, there still remains a challenging open question: how can we learn a robust deep representation that allows for holistic understanding and high-level reasoning from complex data? This CAREER project aims to address this question and is expected to result in novel deep architectures, graphical models, and algorithmic advances for inference, learning, and optimization in deep representation learning. The research outcomes will be disseminated through publications, talks, and tutorials. In addition to advancing the state of the art in deep learning and the many applications it entails, the project will integrate research and education through 1) developing courses in machine learning that include deep learning as a key topic; 2) mentoring significant graduate and undergraduate research activities; and 3) reaching out to K-12 students via hosting demo sessions and mentoring for science fair/research projects. This project investigates the following closely interrelated and complementary thrusts: First, it develops deep learning algorithms to disentangle factors of variation from complex data. This is done by modeling higher-order interactions between multiple groups of latent variables with a deep generative model (e.g., modeling face images via interaction of latent factors that correspond to identity, viewpoint, and emotion). In addition to better generalization, this approach is amenable to high-level reasoning, such as making analogies. Modeling higher-order interaction will be approached by learning a sub-manifold for each factor of variation, where correspondence information is used for regularizing the latent representation. The project will also develop weakly-supervised and semi-supervised disentangling algorithms that automatically establish correspondences without manual supervision. Second, the project develops deep representation learning methods for structured prediction problems. Specifically, it will develop a graphical model with deep representations that can model complex dependencies between output variables. This framework can be also viewed as data-driven modeling of higher-order prior on structured data, and can be used for modeling higher-order conditional random fields that permit efficient inference and learning. In addition, the project develops stochastic conditional generative models for structured prediction problems that involve uncertainty (i.e., one-to-many mappings). Third, the project develops novel deep learning algorithms for constructing shared representations from multiple heterogeneous input modalities, such as image and text, audio and video, and multiple sensor streams. The main idea is to separately model conditional distribution of each input modality given other modalities. This approach addresses the well-known difficulty of modeling a joint distribution across heterogeneous multimodal input, and provides a theoretical analysis on conditions under which the approach can recover a consistent generative model. This formulation allows for robust recognition and high-level reasoning from heterogeneous multimodal data. Overall, these three thrusts are complementary and are expected to play synergistic roles in tackling a broader range of AI problems and moving beyond the current state-of-the-art in deep learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Toward Scalable Life-long Representation Learning
海外基金