EAGER-DynamicData: Principled and Scalable Probabilistic Frameworks for Dynamic Multi-modal Data
EAGER-DynamicData: Principled and Scalable Probabilistic Frameworks for Dynamic Multi-modal Data
批准号:
1462502
负责人:
Lawrence Carin
金额:
$10.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-09-01 至 2017-08-31
中文摘要
大数据现象的出现产生了大量、高度异质和多模式、动态演变以及不完整、嘈杂和不精确的数据收集。这些特征在来自不同领域的数据中变得越来越普遍,例如机器人、认知神经科学、传感器生成的数据(例如,在地球科学和遥感),并在网上动态演变的数据。异构性,复杂性,动态演化,以及通常的实时处理要求,要求的方法都是严格的统计以及计算可扩展性。此外,在 * 测试时间 * 执行快速特征提取和/或预测是另一个关键要求,特别是在涉及高速到达的动态数据的问题中。该项目将在可扩展的统计方法上进行创新,以便从这种海量动态多模态数据中学习,重点是设计用于此类数据的多层潜在特征提取的新型概率模型。这些数据的多层潜在特征表示将有助于捕获潜在的动态,并允许协调由于不同的数据类型和不同模态之间广泛不同的空间和时间分辨率而产生的数据异质性,同时也可用于广泛的基本数据分析任务,例如分类,聚类和预测缺失数据。同时,重点也将放在开发在测试时有效的方法上,以便可以在真实的时间内进行快速特征提取和预测,使这些方法易于应用于动态流数据。EARly探索性研究资助(EAGER)项目致力于超越目前用于这些问题的现有特设方法,并开发一种基于概率的,统计上严格的,和计算可扩展的框架,基于贝叶斯和非参数贝叶斯建模。采用贝叶斯生成建模方法自然能够对数据的动态行为进行建模,并无缝集成各种类型的数据,同时处理数据的缺失、噪声和不精确性等问题。此外,非参数贝叶斯处理将提供急需的建模灵活性,并解决现有深度学习模型的许多局限性,例如,通过消除大量手动调整的需要,结合关于模型参数的丰富先验知识,并允许跨多种数据模态自然共享统计强度。为了处理相关的计算挑战,该框架将以在线贝叶斯推理方法的形式提供新颖的推理机制,该方法将自然地处理动态实时数据,以及并行和分布式贝叶斯推理方法,以处理对于单个计算节点的容量(存储和/或计算)来说太大的大量多模态数据。此外,由于其量化模型不确定性的固有能力,所提出的贝叶斯框架将自然地促进模型计算(推断)和数据获取之间的动态集成,并且帮助设计知情的数据获取(即,“主动”感测)方法。该项目的总体目标是帮助协同机器学习中的两个重要研究方向-非参数贝叶斯方法和深度学习方法。通过设计可扩展的非参数贝叶斯解决方案来解决深度学习方法所应用的问题,该项目将说服深度学习方法的怀疑者更公开地采用这些方法。与此同时,深度学习正在被用于的令人信服的问题和应用范围,将从实际意义上扩大非参数贝叶斯方法的吸引力。我们预计,这两个领域之间的协同作用将大大推进这两个领域的最新技术水平。
英文摘要
Emergence of the Big Data phenomenon has given rise to data collections that are massive, highly heterogeneous and multi-modal, dynamically evolving, as well as incomplete, noisy and imprecise. These characteristics are becoming increasingly prevalent in data from a diverse range of domains, such as robotics, cognitive neuroscience, sensor generated data (e.g., in geoscience and remote sensing), and the dynamically evolving data on the web. The heterogeneity, complexity, dynamic evolution, and the often real-time processing requirements, call for methods that are both statistically rigorous as well as computationally scalable. Moreover, performing fast feature-extraction and/or predictions at *test time* is another key requirement, especially in problems involving dynamic data arriving at high speeds. This project will innovate on scalable statistical methods for learning from such massive dynamic multi-modal data, with a focus on designing novel probabilistic models for multi-layer latent feature extraction for such data. These multi-layer latent feature representations of the data will help capture the underlying dynamics and allow reconciling the data heterogeneity arising due to diverse data types and widely differing spatial and temporal resolutions across the different modalities, while also being useful for a wide range of fundamental data analysis tasks, such as classification, clustering, and predicting missing data. At the same time, the focus will also be on developing methods that are efficient at test time, so that fast feature extraction and predictions can be made in real time, to make these methods readily applicable to dynamic streaming data.This EArly Grant for Exploratory Research (EAGER) project endeavors to move beyond existing ad hoc approaches currently used for these problems, and develop a probabilistically grounded, statistically rigorous, and computationally scalable framework, based on Bayesian and nonparametric Bayesian modeling. Taking a Bayesian generative modeling approach will naturally enable modeling the dynamic behavior of the data and seamlessly integrate diverse types of data, while handling issues such as missingness, noise and the imprecise nature of the data. In addition, the nonparametric Bayesian treatment will provide the much-needed modeling flexibility and address many of the limitations of the existing Deep Learning models, e.g., by doing away with the need of extensive hand-tuning, incorporating rich prior knowledge about the model parameters, and allowing a natural sharing of statistical strength across the multiple data modalities. To handle the associated computational challenges, the framework will provide novel inference machinery in form of online Bayesian inference methods that will naturally handle dynamic, real-time data, and parallel and distributed Bayesian inference methods to handle massive multi-modal data that are too large for the capacity (storage and/or computational) of a single computing node. Furthermore, due to its inherent ability of quantifying model uncertainty, the proposed Bayesian framework will naturally facilitate a dynamic integration between model computation (inference) and data acquisition, and help design informed data acquisition (i.e., "active" sensing) methods in the context of dynamic multi-modal data. An overarching goal of this project is to also help synergize two important research directions in machine learning - nonparametric Bayesian methods and Deep Learning methods. By designing scalable nonparametric Bayesian solutions to the type of problems Deep Learning methods have been applied for, the project will convince the skeptics of Deep Learning methods to adopt these methods more openly. At the same time, the compelling range of problems and applications Deep Learning are being used for, will broaden the appeal of nonparametric Bayesian methods from a practical sense. We expect this synergy between these two areas will significantly advance the state-of-the-art in both areas.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RIA: Short-Pulse, Ultra-Wideband Scattering Range
-
批准号:9596219
-
项目类别:Standard Grant
-
资助金额:$3.0万
-
财政年份:1995
-
负责人:Lawrence Carin
-
依托单位:
RIA: Short-Pulse, Ultra-Wideband Scattering Range
-
批准号:9211353
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:1992
-
负责人:Lawrence Carin
-
依托单位:
海外基金