Developing Statistical Tools for Data integration and Data Fusion for Finite Population Inference
Developing Statistical Tools for Data integration and Data Fusion for Finite Population Inference
批准号:
2242820
负责人:
Jae-Kwang Kim
金额:
$37.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31
中文摘要
该研究项目将为数据集成和数据融合开发统计学习工具。鉴于新数据来源的激增,研究人员越来越多地利用方便但往往不受控制的大数据来源、网络调查面板和管理数据。数据集成是一个新兴的研究领域,它以可靠的方式将多个数据源组合在一起。然而,用于数据整合的统计工具有限。这项研究的结果将对大数据复杂调查数据的分析以及从多个数据集得出的科学结论产生重大影响。研究人员计划与不同统计机构的研究人员积极合作,并将进行应用,以展示新方法在各种环境中的价值。这项研究的结果将通过出版物、演讲、短期课程、网络研讨会和软件进行传播。研究生将参与这项研究。这项研究项目将为数据集成和融合提供统计和机器学习工具。统计机构面临着越来越大的压力,要求它们利用方便但往往不受控制的数据源。虽然这种数据来源为大量变量和人口要素提供了及时的数据,但由于固有的选择偏见,它们往往不能代表感兴趣的目标人口。通过使用独立的概率样本作为校准样本,可以减少方便样本中的选择偏差;然而,用于数据整合的统计工具还不能令人满意。在抽样调查研究中,结合多数据源的统计推断是一个研究相对较少的课题。这项研究将扩大调查数据分析的范围,为数据集成提供许多统计和机器学习工具,并通过实例应用加强数据集成的使用。研究人员将讨论重要的研究主题,如使用现代机器学习工具的大规模归罪,使用信息投影的倾向分数加权,使用高维协变量的校准加权,用于数据集成的多重偏差校准,以及用于数据融合的最佳估计和抽样设计。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This research project will develop statistical learning tools for data integration and data fusion. Given the proliferation of new data sources, researchers are increasingly utilizing convenient but often uncontrolled big data sources, web survey panels, and administrative data. Data integration is an emerging field of study that combines multiple data sources in a reliable way. However, statistical tools for data integration are limited. The results of this research will significantly impact the analysis of complex survey data with big data, as well as scientific conclusions drawn from multiple data sets. The investigator plans to actively collaborate with researchers at different statistical agencies, and applications will be conducted to demonstrate the value of the new methods in various settings. The results of this research will be disseminated via publications, presentations, short courses, webinars, and software. Graduate students will be involved in the conduct of the research.This research project will produce statistical and machine learning tools for data integration and fusion. Statistical agencies face increasing pressure to utilize convenient but often uncontrolled sources of data. While such data sources provide timely data for a large number of variables and population elements, they often fail to represent the target population of interest because of inherent selection biases. By using an independent probability sample as a calibration sample, the selection bias in the convenience sample can be reduced; however, the statistical tools for data integration are not yet satisfactory. In survey sampling research, statistical inference combining multiple data sources is a relatively understudied topic. This research will expand the scope of survey data analysis by providing numerous statistical and machine learning tools for data integration and by enhancing the use of data integration through example applications. The investigator will address important research topics such as mass imputation using modern machine learning tools, propensity score weighting using information projection, calibration weighting with high dimensional covariates, multiple bias calibration for data integration, and optimal estimation and sampling design for data fusion.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Innovations in Statistical Methodology for Complex Surveys
-
批准号:1733572
-
项目类别:Standard Grant
-
资助金额:$43.0万
-
财政年份:2017
-
负责人:Jae-Kwang Kim
-
依托单位:
Fractional Imputation for Incomplete Data Analysis
-
批准号:1324922
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2013
-
负责人:Jae-Kwang Kim
-
依托单位:
海外基金