Semiparametric Efficient and Robust Inference on High-Dimensional Data
Semiparametric Efficient and Robust Inference on High-Dimensional Data
批准号:
2310578
负责人:
Alexander Giessing
金额:
$17.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-15 至 2026-05-31
中文摘要
为科学或工业目的收集的大多数数据都存在缺失,缺乏关于某些特征的信息,这些特征在采样阶段由于实验设计、不符合规定或技术问题等因素而丢失。这个问题在高维数据集中尤其普遍,其中每个观测值包含许多特征。如果不能有效地解决数据缺失和不完整的问题,就会导致低效和有偏见的推理。此外,当最初设计用于低维不完整数据的传统统计方法应用于高维数据时,它们可能产生误导性的科学发现。因此,迫切需要创新的统计方法,专门针对高维不完整数据的推理。例如,在癌症研究中,下一代测序技术允许进行全面的基因组分析。然而,由于技术限制和肿瘤异质性,这种分析通常是不完整的。在推断不完整的特征后进行有效推断的能力对推进癌症治疗具有重要意义。本研究项目包括:(1)开发高维不完整和缺失数据的推理方法;(2)通过出版物、研讨会和公开发布软件向统计界传播所得技术;(3)培养高维统计与概率论方面的博士生;(4)通过介绍性阅读小组、会议演讲和其他外展活动,增加高中生、本科生和代表性不足群体成员对统计和概率论的接触。该项目将提供广泛的指导、教育和专业发展机会,以培训不同职业水平的下一代统计学家和数据科学家。提出的研究有两个主要目标,即(1)在调整高维不完整数据后,开发一种半参数有效的推理方法;(2)开发在不完整数据下保持正确大小和功率的高维假设的单样本和双样本自举检验。不完整数据的推断需要通过逆概率加权或单/多重输入进行仔细调整。这些调整中的任何估计错误都会传播到后续的分析中。为了应对这一挑战并实现第一个目标,该项目将开发综合推理和调整程序,将归算/重新加权视为推理过程的一个组成部分,而不是一个单独的麻烦步骤。将提供几个典型的不完整数据问题的具体解决方案。对于第二个目标,主要挑战在于设计一个能够准确地解释输入/重新加权的可变性的自举程序。为了实现这一目标,该项目将开发一种新的参数化高维引导程序,可以利用这些信息。将提供离散/分类和连续数据的不同自举测试。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Most data collected for scientific or industrial purposes suffer from missingness, lacking information about certain features that were lost at the sampling stage due to factors such as experimental design, non-compliance, or technical problems. This issue is particularly prevalent in high-dimensional data sets, where each observation comprises many features. Failing to effectively address the issue of missing and incomplete data results in inefficient and biased inference. Moreover, when traditional statistical methods, originally designed for low-dimensional incomplete data, are applied to high-dimensional data, they can yield misleading scientific discoveries. Therefore, there is an urgent need for innovative statistical methodologies specifically tailored to inference on incomplete data in high dimensions. For instance, in cancer research, next-generation sequencing technology allows for comprehensive genomic profiling. However, due to technical limitations and tumor heterogeneity, this profiling is typically incomplete. The ability to conduct valid inference after imputing incomplete profiles holds significant implications for advancing cancer treatment. This research project involves (1) developing methodology for inference on high-dimensional incomplete and missing data; (2) disseminating the resulting techniques to the statistics community via publications, seminars, and public release of software; (3) training PhD students in high-dimensional statistics and probability theory; and (4) increasing the exposure of high school students, undergraduates, and members of underrepresented groups to statistics and probability theory via introductory reading groups, conference presentations, and other outreach activities. The project will provide a broad range of mentoring, educational and professional development opportunities to train the next generation of statisticians and data scientists at various career levels.The proposed research has two main objectives, namely (1) to develop a semiparametric efficient approach to inference after adjusting for incomplete data in high dimensions, and (2) to develop one- and two-sample bootstrap tests for high-dimensional hypotheses that retain correct size and power under incomplete data. Inference with incomplete data requires careful adjustments through inverse probability weighting or single/ multiple imputation. Any estimation error in these adjustments propagates into subsequent analyses. To address this challenge and achieve the first objective, the project will develop combined inference and adjustment procedures which treat imputation/ re-weighting not as a separate nuisance step, but as an integral part of the inference process. Specific solutions to several canonical incomplete data problems will be provided. For the second objective, the main challenge lies in designing a bootstrap procedure that accurately accounts for the variability of imputation/ re-weighting. To meet this objective, the project will develop a new parametric high-dimensional bootstrap procedure that can leverage such information. Different bootstrap tests for discrete/ categorical and continuous data will be provided.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金