课题基金 / 基金详情

Statistical Design, Sampling, and Analysis for Large Scale Experiments

Statistical Design, Sampling, and Analysis for Large Scale Experiments
大规模实验的统计设计、采样和分析
批准号:
1916467
负责人:
Lulu Kang
金额:
$12.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-09-01 至 2023-08-31

项目摘要

项目成果

Lulu Kang的其他基金

相似基金

相关文献

中文摘要
翻译
在大数据范式中,即使是受控实验也可以变得大规模,即样本量巨大,输入变量的维数很高。这种“大数据”问题挑战了许多统计方法,显著增加了估计和推理的计算量。在这个项目中,PI着重于大规模实验的具体实例,并开发了一套关于实验设计、抽样和分析的新理论和方法。这项研究有两个主要部分。在第1部分中,PI关注的是包含大量协变量的实验类型。例如,在临床试验中,协变量可以是患者丰富的病史。如何为每位患者分配治疗环境?PI通过一般的实验设计框架提供答案,以便尽管协变量的影响,也能准确地估计治疗效果。在第2部分中,PI将重点介绍高斯过程(GP)回归,这是最流行的统计学习工具之一。所需的计算量对于分析气候模式模拟等大规模实验来说是令人望而却步的。PI开发了一个降维框架和一种主动学习方法,显著提高了GP模型的效率和准确性。这里考虑了三种主要的方法。在第1-3部分中,PI引入了一种新的基于差异的设计,以实现具有大维度协变量的实验的协变量平衡。差异准则还具有吸引人的理论性质,可以更准确地估计参数,包括治疗效果和协变量的影响。提出了离线实验和在线实验的优化设计算法。在第4部分中,PI开发了一种新的降维方法,用于寻找GP模型的低维核函数的最优凸组合。结果表明,所提出的方法对某些类型的底层函数的近似计算量大大减少,而且精度更高。第五部分提出了一种基于广义库克距离的GP回归主动学习方法。它比标准的随机抽样方法更有效。该研究思路新颖、理论严谨、实践实用,将为实验统计设计与分析开辟新的方向。PI有一个详细的教育计划,以这个项目的研究成果为基础,开发新的课程模块、教程和讲习班。研究成果易于应用于需要大规模数据收集和分析的各种科学、工程、医学等领域。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In the big data paradigm, even the controlled experiments can become large-scale, in the sense that the sample size is massive, and the dimension of the input variables is high. Such "big data" problem challenges many statistical approaches and significantly increases the amount of computation in estimation and inference. In this project, the PI focuses on specific instances of large-scale experiments and develops a set of novel theories and methodologies on experimental design, sampling, and analysis. The research has two major parts. In Part 1, the PI focuses on the type of experiments that contains a large dimension of covariate variables. For example, in a clinical trial, the covariates can be patients' rich medical history. How should the treatment settings be assigned to each patient? The PI provides the answer through a general experimental design framework so that the treatment effects are estimated accurately despite the influence of the covariates. In Part 2, the PI focuses on the Gaussian Process (GP) regression, one of the most popular statistical learning tools. The computation required is prohibitive for analyzing large-scale experiments such as the climate model simulations. The PI develops a dimension reduction framework and an active learning method that significantly improves the efficiency and accuracy of the GP model.Three major methodologies are considered. In Parts 1-3, the PI introduces a new discrepancy-based design to achieve covariate balance for experiments with a large dimension of covariates. The discrepancy criterion also has appealing theoretical properties that lead to a more accurate estimation of the parameters including both treatment effects and covariates' effects. Optimal design algorithms are developed for both offline and online experiments. In Part 4, the PI develops a novel dimension reduction method that finds the optimal convex combination of low-dimension kernel functions for the GP model. It is shown that the proposed method is a significantly less computational and more accurate approximation of certain types of underlying functions. In Part 5, an active learning method based on the generalized Cook's Distance is developed for the GP regression. It is more efficient than the standard random sampling method. The research is novel in ideas, rigorous in theories, and useful in practice, and will open new directions in the statistical design and analysis of experiments area. The PI has a detailed education plan to develop new course modules, tutorials, and workshops based on the research products from this project. The research outcomes are readily applicable to a variety of scientific, engineering, medicine and other fields where large-scale data collection and analysis are demanded.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
Locally Optimal Design for A/B Tests in the Presence of Covariates and Network Dependence
存在协变量和网络依赖性的情况下 A/B 测试的局部最优设计
DOI: 10.1080/00401706.2022.2046169
发表时间: 2022
期刊: Technometrics
影响因子: 2.5
作者: [Zhang, Qiong, Kang, Lulu]
通讯作者: Kang, Lulu
Bayesian D-Optimal Design of Experiments with Quantitative and Qualitative Responses
具有定量和定性响应的贝叶斯 D 优化实验设计
DOI: 10.51387/23-nejsds30
发表时间: 2023
期刊: The New England Journal of Statistics in Data Science
影响因子: --
作者: [Kang, Lulu, Deng, Xinwei, Jin, Ran]
通讯作者: Jin, Ran
DOI: 10.1080/00401706.2020.1817790
发表时间: 2020-10-12
期刊: TECHNOMETRICS
影响因子: 2.5
作者: [Chen, Jiuhai, Kang, Lulu, Lin, Guang]
通讯作者: Lin, Guang
A Maximin Φp-Efficient Design for Multivariate Generalized Linear Models
多元广义线性模型的最大最小Ψ有效设计
DOI: 10.5705/ss.202020.0278
发表时间: 2023
期刊: Statistica Sinica
影响因子: 1.4
作者: [Li, Yiou, Kang, Lulu, Deng, Xinwei]
通讯作者: Deng, Xinwei
共 7 条
    Energetic Variational Inference: Foundations, Algorithms, and Applications
    • 批准号:
      2153029
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $30.0万
    • 财政年份:
      2022
    • 负责人:
      Lulu Kang
    • 依托单位:
    Collaborative Research: Experimental Design and Analysis of Quantitative-Qualitative Responses in Manufacturing and Biomedical Systems
    • 批准号:
      1435902
    • 项目类别:
      Standard Grant
    • 资助金额:
      $11.79万
    • 财政年份:
      2014
    • 负责人:
      Lulu Kang
    • 依托单位:
    国内基金
    海外基金
    Applications of AI in Market Design
    • 批准号:
      --
    • 项目类别:
      外国青年学者研 究基金项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
      Manshu Khanna
    • 依托单位:
    基于“Design-Build-Test”循环策略的新型紫色杆菌素组合生物合成研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2021
    • 负责人:
    • 依托单位:
    在噪声和约束条件下的unitary design的理论研究
    • 批准号:
      12147123
    • 项目类别:
      专项基金项目
    • 资助金额:
      18万元
    • 批准年份:
      2021
    • 负责人:
      顾炎武
    • 依托单位: