Representation and Subspace Learning for Decentralized and Dependent Data
Representation and Subspace Learning for Decentralized and Dependent Data
批准号:
2015366
负责人:
Ziwei Zhu
金额:
$19.9万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-01 至 2022-09-30
中文摘要
学习高维数据的简明和信息表示是现代数据分析成功的先决条件。然而,近年来出现了许多非标准的数据机制,给表示学习带来了前所未有的挑战。第一种情况是数据是分散的,也就是说,它们分散在不同的地方,在这些地方之间的通信受到高度限制。这对于在全球范围内收集数据的跨国公司来说很常见,但由于网络带宽或法律政策的限制,它们无法汇总数据。第二种情况是数据表现出明显的时间依赖性,如股票价格、交通流量和临床试验。该项目将开发具有理论保证的新颖统计方法来处理这些现代数据制度。它还旨在在这些重要的问题设置下培训下一代数据科学家。首席研究员(PI)将为分散和依赖数据的子空间和表示学习开发新的方法和理论。对于分散的数据,PI计划设计和研究一种新的方法框架,用于一般潜在变量模型的分布式估计。该框架只需要一轮模型参数的通信,适应范围广泛的复杂潜变量模型(包括基于深度神经网络的模型),并已被证明比现有方法产生更好的数值性能。PI将考虑的另一个更具体的设置是奇异空间的分布估计,并将其应用于谱聚类。对于相关数据,PI将专注于学习低秩马尔可夫转移核的顶部奇异空间来执行状态压缩和降维。PI计划通过核范数惩罚或等级约束最大化对数似然来解决问题。本文将明确地推导出m估计量的统计率,并开发新的优化算法来计算这些具有收敛性保证的问题。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Learning concise and informative representations of high-dimensional data is a precursor to the success of modern data analytics. However, recent years have witnessed many non-standard data regimes that impose unprecedented challenges for representation learning. The first scenario is that data are decentralized, that is, they are scattered across different places across which the communication is highly restricted. This is common for international companies that collect data worldwide, but cannot aggregate them due to constraints on network bandwidth or legal policies. The second scenario is that data exhibit significant temporal dependence, as seen in stock prices, traffic flow, and clinical trials. This project will develop novel statistical methods with theoretical guarantees to handle these modern data regimes. It also aims to train the next generation of data scientists under these important problem setups. The principal investigator (PI) will develop novel methods and theory for subspace and representation learning for decentralized and dependent data. For decentralized data, the PI plans to design and study a new methodological framework for distributed estimation of a general latent variable model. This framework requires only one round of communication of model parameters, adapts to a wide range of complex latent variable models (including those based on deep neural nets) and has been shown to yield superior numerical performance over existing approaches. Another more specific setup that the PI will consider is distributed estimation of singular spaces, with applications to spectral clustering. For dependent data, the PI will focus on learning the top singular space of a low-rank Markov transition kernel to perform state compression and dimension reduction. The PI plans to solve the problem via maximizing the log-likelihood with either nuclear-norm penalty or rank constraint. The statistical rate of the resulting M-estimator will be explicitly derived, and new optimization algorithms will be developed to compute these problems with convergence guarantee.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
海外基金