课题基金 / 基金详情

高维非独立同分布数据的分布式推断

批准号:
12101240
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
王小舟
依托单位:
学科分类:
统计推断与统计计算
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
王小舟

项目摘要

结项摘要

相似基金

相关文献

中文摘要
分布式推断是统计学热门研究问题,分布式算法不仅具有计算高效性,而且与经典方法具有渐近一致的统计效率,因此可以有效应对海量数据带来的技术挑战。现有的分布式研究主要基于数据的独立同分布假设,然而在实际应用中,不同局部机器收集得到的数据可能存在机器差异性,即来自不同的子总体和环境。非独立同分布数据会导致经典的分布式算法及对应理论推断失效,给当前的分布式研究带来很大的挑战。本项目拟针对高维非独立同分布数据,研究分布参数估计问题。通过引入局部机器特有参数,建立分布调整机制,并构造通信高效的分布式算法。进一步探索稀疏回归和支持向量机模型下的分布式算法构造,根据非独立同分布数据结构结合正则化方法研究高维模型下的参数估计和统计推断问题。所得算法将达到近似最优统计误差率,并被应用于生物、医学、工程学等领域解决实际问题。
英文摘要
Distributed inference is a popular topic in statistics. Distributed algorithms not only are computationally efficient but also achieve the optimal statistical efficiency, hence they can effectively overcome the challenges brought by the growing size of modern data. Most existing distributed methods focus on independent and identically distributed (i.i.d.) data, while in practice, data collected from different local machines may be machine-specific, i.e., coming from various sub-populations and environments. Non-i.i.d. data make the classical distributed methodologies and theories invalid, which brings huge challenges to existing distributed studies. In this project, we propose to study the parameter estimation of distributions for high-dimensional non-i.i.d. data. Through introducing machine-specific parameters, we establish an adjustment mechanism for distributions and construct communication-efficient distributed algorithms. We further explore distributed algorithms for sparse regression and support vector machine, conduct parameter estimation and statistical inference based on non-i.i.d. data with regularization methods for high-dimensional models. The resulting methods will attain near-optimal statistical error rates and be evaluated by practice applications in biology, medicine, engineering and other fields.
分布式推断是统计学的热门研究问题,相关方法可以有效应对海量数据带来的技术挑战。现有的分布式算法主要针对独立同分布数据,假设所有局部机器中的数据均服从相同分布。然而在实际应用中,不同局部机器收集得到的数据可能来自不同的子总体和环境,部分机器和数据甚至可能受到攻击或污染的影响,这会导致基于独立同分布数据的分布式算法与理论结果失效。本项目针对分类、回归等统计问题,探究对应的参数估计和统计推断,以此为基础将方法逐步推广至分布式学习与非独立同分布数据下的分布式学习框架。项目从算法构造、理论分析、实际应用等方面均拓展了分布式学习方向,得到的算法与理论结果可广泛应用于处理各类大规模数据集,用来解决医学、生物学、经济学、工程学等领域的实际问题。项目成果具有重要的科学技术意义和广泛的应用场景,也为进一步的研究打下了坚实的基础。
国内基金
海外基金