一类最优的子抽样设计及其统计分析方法
批准号:
12071088
项目类别:
面上项目
资助金额:
51.0 万元
负责人:
郁文
依托单位:
学科分类:
贝叶斯统计与统计应用
结题年份:
2024
批准年份:
2020
项目状态:
已结题
项目参与者:
郁文
中文摘要
大数据时代,数据体量大幅增加,分析方法的计算复杂度显著提高,导致数据采集耗费的经济成本与日俱增,数据分析对于计算机存储与运算能力的要求也越来越高。当抽样成本与计算资源有限,或者对于数据分析的时效性要求较高时,子抽样方法可以有效地降低待分析数据体量,从而缩短分析时间,节省经济成本与计算资源。然而,子抽样因只使用原始数据的一部分,可能会影响数据分析方法的正确性与效率,因此,如何尽量优选数据、保证分析方法正确、减少效率损失是子抽样设计的重点。本项目提出一类新的子抽样设计思想,可以根据不同的数据分析目标对应地给出最优抽样设计,适用于多种常用的模型框架与不同的数据类型,对模型错误假定具备稳健性。除抽样设计外,本项目结合逆概率加权技术提出基于子抽样数据的统计分析方法,抽样设计的最优性使得新抽样方法在相同的抽样成本下渐近地达到最优效率。本项目的研究将丰富大规模数据处理的抽样与分析手段。
英文摘要
In the era of big data, with the rapid growth of data volume and significant raise of computation complexity, the economic cost of data collection is increasing rapidly and the requirement of computer storage and computation ability for data analysis is also getting higher and higher. When there is limited sampling cost and computation capacity, or when there is requirement for the speed of data analysis, sumsampling can help to reduce the data volume significantly, resulting in reduction in computational time and cost on data collection. However, since subsampling only use a fraction of the original data, it may affect the validity and efficiency of the analysis methods. Thus, it is quite important to consider how to optimally selection the data, to keep the validity of the methods, and to reduce the efficiency lose. In this project, we propose an optimal subsampling design which can adaptively give out optimal design according to different purpose of data analysis. The proposed design can be used in various model frameworks and different kinds of data. It possesses robustness against model misspecification. Beside the design itself, by incorporating the idea of inverse probability weighting, we propose the related statistical methods based on the subsample data. The optimality of the proposed method helps the design to reach the optimal efficiency asymptotically under the same sampling cost. The research of this project will enrich the sampling and data analysis techniques in dealing with large scale data.
大数据时代,数据体量的增加和算法复杂度的提高,使得数据处理的计算与经济成本与日俱增。当抽样成本与计算资源有限,或者对数据分析时效性要求较高时,子抽样方法可以有效地降低待分析数据体量,从而缩短分析时间,节省计算成本。然而,子抽样方法减少了数据使用量,会影响数据分析的精确性与效率。因此需要在计算成本与分析效率之间达到平衡。本项目提出一类新的子抽样设计思想,可在多种模型假设下根据不同的数据分析目标对应地给出最优子抽样设计,其最优性体现为在相同的抽样成本下渐近地达到最优效率,同时对模型的错误假定具备一定的稳健性。. 在本项目研究期间,项目负责人与团队成员依照项目计划开展研究工作,在一系列相关的研究课题中取得进展与成果。研究方面,具体成果如下。首先,项目团队在广义线性模型假设下,基于逆概率加权估计的渐近性质推导最优子抽样设计方案,给出相应的理论性质与数值实验结果,验证抽样方案的有效性。相关成果发表于国际统计学术期刊Electronic Journal of Statistics上。其次,项目团队在Cox模型假设下,基于逆概率加权估计的渐近性质推导右删失数据的最优子抽样设计方案,给出相应的理论性质与数值实验结果,验证抽样方案的有效性。相关成果发表于国际统计学术期刊Journal of Applied Statistics上。此外,项目团队在项目资金的支持下开展一系列相关的研究工作,包括带有异常值的点过程聚类、有偏抽样下二分类数据半监督学习、双侧截断数据变量选择、基于深度神经网络危险率模型拟合以及疫情防控措施等,相关成果发表在国际会议NeurIPS与国内学术期刊Fundamental Research、中国科学、应用概率统计上。人才培养方面,项目资金辅助培养博士研究生5人、硕士研究生1人。项目成员也多次参加国内外学术会议,报告研究成果,并与国内外同行进行了密切的互动与交流。
终端抽样设计及其在生存分析中的应用
-
批准号:11671097
-
项目类别:面上项目
-
资助金额:48.0万元
-
批准年份:2016
-
负责人:郁文
-
依托单位:
基于经验似然方法的保序推断
-
批准号:11101091
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:郁文
-
依托单位:
国内基金
海外基金