课题基金 / 基金详情

CAREER: New Directions in Deep Representation Learning from Complex Multimodal Data

CAREER: New Directions in Deep Representation Learning from Complex Multimodal Data
职业:复杂多模态数据深度表示学习的新方向
批准号:
1453651
负责人:
Honglak Lee
金额:
$48.86万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2023-08-31

项目摘要

项目成果

Honglak Lee的其他基金

相似基金

相关文献

中文摘要
翻译
深度学习的目标是学习具有层次结构和组合结构的数据的抽象表示。深度学习方法可以有效地从高维输入数据中学习判别特征(例如,用于分类),并已成功地应用于许多现实世界的问题,如图像分类、语音识别和文本建模。尽管取得了这些成功,但仍然存在一个具有挑战性的开放性问题:我们如何才能学习一个强大的深度表征,从而从复杂的数据中进行整体理解和高级推理?这个CAREER项目旨在解决这个问题,并有望为深度表示学习中的推理、学习和优化带来新颖的深度架构、图形模型和算法进步。研究成果将通过出版物,讲座和教程传播。除了推进深度学习的最新技术及其所需的许多应用之外,该项目还将通过以下方式整合研究和教育:1)开发机器学习课程,将深度学习作为一个关键主题;2)指导重要的研究生和本科生研究活动;3)通过主持演示会议和指导科学展览/研究项目来接触K-12学生。该项目研究了以下密切相关和互补的重点:首先,它开发了深度学习算法,以从复杂数据中分离出变化因素。这是通过使用深度生成模型对多组潜在变量之间的高阶相互作用进行建模来完成的(例如,通过与身份、观点和情感相对应的潜在因素的相互作用对人脸图像进行建模)。除了更好的泛化之外,这种方法还适用于高级推理,例如进行类比。建模高阶交互将通过学习每个变化因素的子流形来接近,其中对应信息用于正则化潜在表示。该项目还将开发弱监督和半监督解纠缠算法,这些算法可以在没有人工监督的情况下自动建立对应关系。其次,该项目开发了结构化预测问题的深度表示学习方法。具体来说,它将开发一个具有深度表示的图形模型,可以对输出变量之间的复杂依赖关系进行建模。这个框架也可以看作是结构化数据上的高阶先验的数据驱动建模,并且可以用于建模允许有效推理和学习的高阶条件随机场。此外,该项目还为涉及不确定性(即一对多映射)的结构化预测问题开发了随机条件生成模型。第三,该项目开发了新的深度学习算法,用于从多种异构输入模式(如图像和文本、音频和视频以及多个传感器流)构建共享表示。主要思想是在给定其他模态的情况下,分别对每种输入模态的条件分布进行建模。该方法解决了跨异构多模态输入的联合分布建模的众所周知的困难,并提供了对该方法可以恢复一致生成模型的条件的理论分析。该公式允许对异构多模态数据进行鲁棒识别和高级推理。总的来说,这三个重点是互补的,预计将在解决更广泛的人工智能问题和超越当前深度学习的最新技术方面发挥协同作用。
英文摘要
The goal of deep learning is to learn an abstract representation of data with a hierarchical and compositional structure. Deep learning methods can effectively learn discriminative features from high-dimensional input data (e.g., for classification), and have been successfully applied to many real-world problems, such as image classification, speech recognition, and text modeling. Despite these successes, there still remains a challenging open question: how can we learn a robust deep representation that allows for holistic understanding and high-level reasoning from complex data? This CAREER project aims to address this question and is expected to result in novel deep architectures, graphical models, and algorithmic advances for inference, learning, and optimization in deep representation learning. The research outcomes will be disseminated through publications, talks, and tutorials. In addition to advancing the state of the art in deep learning and the many applications it entails, the project will integrate research and education through 1) developing courses in machine learning that include deep learning as a key topic; 2) mentoring significant graduate and undergraduate research activities; and 3) reaching out to K-12 students via hosting demo sessions and mentoring for science fair/research projects. This project investigates the following closely interrelated and complementary thrusts: First, it develops deep learning algorithms to disentangle factors of variation from complex data. This is done by modeling higher-order interactions between multiple groups of latent variables with a deep generative model (e.g., modeling face images via interaction of latent factors that correspond to identity, viewpoint, and emotion). In addition to better generalization, this approach is amenable to high-level reasoning, such as making analogies. Modeling higher-order interaction will be approached by learning a sub-manifold for each factor of variation, where correspondence information is used for regularizing the latent representation. The project will also develop weakly-supervised and semi-supervised disentangling algorithms that automatically establish correspondences without manual supervision. Second, the project develops deep representation learning methods for structured prediction problems. Specifically, it will develop a graphical model with deep representations that can model complex dependencies between output variables. This framework can be also viewed as data-driven modeling of higher-order prior on structured data, and can be used for modeling higher-order conditional random fields that permit efficient inference and learning. In addition, the project develops stochastic conditional generative models for structured prediction problems that involve uncertainty (i.e., one-to-many mappings). Third, the project develops novel deep learning algorithms for constructing shared representations from multiple heterogeneous input modalities, such as image and text, audio and video, and multiple sensor streams. The main idea is to separately model conditional distribution of each input modality given other modalities. This approach addresses the well-known difficulty of modeling a joint distribution across heterogeneous multimodal input, and provides a theoretical analysis on conditions under which the approach can recover a consistent generative model. This formulation allows for robust recognition and high-level reasoning from heterogeneous multimodal data. Overall, these three thrusts are complementary and are expected to play synergistic roles in tackling a broader range of AI problems and moving beyond the current state-of-the-art in deep learning.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Toward Scalable Life-long Representation Learning
海外基金