含有潜在因子的大规模数据统计分析及应用
批准号:
11601501
项目类别:
青年科学基金项目
资助金额:
19.0 万元
负责人:
郑泽敏
依托单位:
学科分类:
贝叶斯统计与统计应用
结题年份:
2019
批准年份:
2016
项目状态:
已结题
项目参与者:
夏婉婉、陈鹏展、孙钟倩
中文摘要
大规模数据的统计推断是当前统计学研究的前沿领域之一。现有的方法通常假设模型中的影响因子均是可观测的。然而实际模型中却可能隐藏有不可观测的潜在因子。本课题以高维统计推断方法的研究为基础,从潜在因子对可观测变量及因变量的影响入手,探索可以对可观测变量及潜在因子作用系数进行推断的统计方法,并寻求针对样本海量增长时含有潜在因子的大规模数据的新型统计分析框架。主要研究工作将从高维潜在因子的估计、高维可观测变量及潜在因子作用系数的统计推断、海量样本情形下的综合统计量三个方面展开。相关研究成果将应用于经济、金融、生物医疗以及机器学习等领域的大规模数据处理。
英文摘要
Statistical inference for large scale data sets is one of the frontier issues in the current studies on statistical science. Most of existing methods assume implicitly that all features in a model are observable. Yet some latent factors may potentially exist in the hidden structure of the original model. Based on the studies on high-dimensional statistical inference, the current project aims to explore statistical methods that can provide inference for the effects of both the observable covariates and latent factors from the impacts of latent features on the observable covariates and outcome. Moreover, we also want to establish new statistical frameworks for large scale data with latent factors when the sample sizes are extraordinarily large. Main research contents of the project include three aspects: estimation of high-dimensional latent factors, statistical inference for the effects of both observable covariates and latent features, and the overall statistics for extraordinarily large sample sizes. The research results will have broad applications in the areas of economics, finance, biomedical treatment and machine learning for large scale data analysis.
大规模数据的统计推断是当前统计学研究的前沿领域之一。现有的方法通常假设模型中的影响因子均是可观测的。然而实际模型中却可能隐藏有不可观测的潜在因子。本课题以高维统计推断方法的研究为基础,从潜在因子对可观测变量及因变量的影响入手,探索可以对可观测变量及潜在因子作用系数进行推断的统计方法,并寻求针对样本海量增长时含有潜在因子的大规模数据的新型统计分析框架。主要研究工作从高维潜在因子的估计、高维可观测变量及潜在因子作用系数的统计推断、海量样本情形下的综合统计量三个方面展开。相关研究成果发表(含接收待发表)于Journal of Machine Learning Research、Operations Research、Statistica Sinica、Computational Statistics and Data Analysis等国际统计学、机器学习及管理优化著名期刊上,将应用于经济、金融、生物医疗以及机器学习等领域的大规模数据处理。
期刊论文列表
专著列表
科研奖励列表
会议论文列表
专利列表
登录
查看更多内容
Scalable Inference for Massive Data
海量数据的可扩展推理
DOI:
--
发表时间:
2017
期刊:
Procedia Computer Science
影响因子:
--
作者:
[Zemin Zheng, Jiarui Zhang, Yinfei Kong, Yaohua Wu]
通讯作者:
Yaohua Wu
Partitioned Approach for High-dimensional Condence Intervals with Large Split Sizes
具有大分割尺寸的高维稠密区间的分区方法
DOI:
--
发表时间:
--
期刊:
Statistica Sinica
影响因子:
1.4
作者:
[Zemin Zheng, Jiarui Zhang, Yang Li, Yaohua Wu]
通讯作者:
Yaohua Wu
DOI:
10.1016/j.csda.2018.04.009
发表时间:
2018
期刊:
COMPUTATIONAL STATISTICS & DATA ANALYSIS
影响因子:
1.8
作者:
[Zheng Zemin, Li Yang, Yu Chongxiu, Li Gaorong]
通讯作者:
Li Gaorong
Scalable Interpretable Multi-Response Regression via SEED
通过 SEED 进行可扩展、可解释的多响应回归
DOI:
--
发表时间:
2016-08
期刊:
Journal of Machine Learning Research
影响因子:
6
作者:
[Zemin Zheng, M. Taha Bahadori, Yan Liu, Jinchi Lv]
通讯作者:
Jinchi Lv
DOI:
--
发表时间:
2017
期刊:
Procedia Computer Science
影响因子:
--
作者:
[Zemin Zheng, Jia Zhou, Xiao Guo, Daoji Li]
通讯作者:
Daoji Li
共 6 条
高维复杂数据分析中具有可重复性的统计学习方法研究及其应用
-
批准号:72071187
-
项目类别:面上项目
-
资助金额:48.0万元
-
批准年份:2020
-
负责人:郑泽敏
-
依托单位:
国内基金
海外基金