Representation and Subspace Learning for Decentralized and Dependent Data
Representation and Subspace Learning for Decentralized and Dependent Data
批准号:
2015366
负责人:
Ziwei Zhu
金额:
$19.9万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-07-01 至 2022-09-30
中文摘要
学习高维数据的简洁和信息表示是现代数据分析成功的前兆。然而,近年来出现了许多非标准数据机制,这给表征学习带来了前所未有的挑战。第一种情况是数据是分散的,也就是说,它们分散在通信受到高度限制的不同地方。这对于在全球范围内收集数据的国际公司来说很常见,但由于网络带宽或法律的政策的限制,这些公司无法将其汇总。第二种情况是数据表现出显著的时间依赖性,如股票价格、交通流量和临床试验。该项目将开发具有理论保证的新统计方法,以处理这些现代数据体系。它还旨在根据这些重要的问题设置培训下一代数据科学家。主要研究者(PI)将为分散和依赖数据的子空间和表示学习开发新的方法和理论。对于分散数据,PI计划设计和研究一种新的方法框架,用于一般潜变量模型的分布式估计。该框架只需要一轮模型参数的通信,适用于各种复杂的潜变量模型(包括基于深度神经网络的模型),并且已被证明比现有方法具有上级数值性能。PI将考虑的另一个更具体的设置是奇异空间的分布式估计,并应用于谱聚类。对于相关数据,PI将专注于学习低秩马尔可夫转移核的顶部奇异空间,以执行状态压缩和降维。PI计划通过最大化具有核范数惩罚或秩约束的对数似然来解决问题。由此产生的M-估计的统计率将显式导出,并将开发新的优化算法来计算这些问题的收敛性guarantee.This奖项反映了NSF的法定使命,并已被认为是值得通过使用基金会的智力价值和更广泛的影响审查标准进行评估的支持。
英文摘要
Learning concise and informative representations of high-dimensional data is a precursor to the success of modern data analytics. However, recent years have witnessed many non-standard data regimes that impose unprecedented challenges for representation learning. The first scenario is that data are decentralized, that is, they are scattered across different places across which the communication is highly restricted. This is common for international companies that collect data worldwide, but cannot aggregate them due to constraints on network bandwidth or legal policies. The second scenario is that data exhibit significant temporal dependence, as seen in stock prices, traffic flow, and clinical trials. This project will develop novel statistical methods with theoretical guarantees to handle these modern data regimes. It also aims to train the next generation of data scientists under these important problem setups. The principal investigator (PI) will develop novel methods and theory for subspace and representation learning for decentralized and dependent data. For decentralized data, the PI plans to design and study a new methodological framework for distributed estimation of a general latent variable model. This framework requires only one round of communication of model parameters, adapts to a wide range of complex latent variable models (including those based on deep neural nets) and has been shown to yield superior numerical performance over existing approaches. Another more specific setup that the PI will consider is distributed estimation of singular spaces, with applications to spectral clustering. For dependent data, the PI will focus on learning the top singular space of a low-rank Markov transition kernel to perform state compression and dimension reduction. The PI plans to solve the problem via maximizing the log-likelihood with either nuclear-norm penalty or rank constraint. The statistical rate of the resulting M-estimator will be explicitly derived, and new optimization algorithms will be developed to compute these problems with convergence guarantee.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
海外基金