Robust and Efficient Statistical Inference in Large Scale Semi-Supervised Settings
Robust and Efficient Statistical Inference in Large Scale Semi-Supervised Settings
批准号:
2113768
负责人:
Abhishek Chakrabortty
金额:
$17.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-01 至 2024-07-31
中文摘要
该项目将开发在半监督设置中进行稳健统计推断的方法。与更传统的数据设置不同,半监督设置的特点是两种类型的可用数据:1)典型的小型或中等规模的标记(或监督)数据,其中包含响应(或结果)和一组协变量(或预测因子)的观察结果,以及2i)更大规模的未标记(或无监督)数据,仅包含对协变量的观察结果。当协变量很容易用于大型队列时,这种设置自然出现,而响应可能由于实际限制而难以获得和/或昂贵。这些在大数据时代的现代研究中越来越重要,因为大型未标记数据库(通常是电子记录)在标记数据之上变得容易获得(和易于处理)。这样的例子在许多学科中无处不在,包括计算机科学、机器学习、计量经济学和电子健康记录和综合基因组学等生物医学应用。因此,在半监督环境中的统计推断具有重要意义。这里的最终问题是调查何时以及如何使用来自大量未标记数据的额外信息来“改进”相应的监督方法,其中改进可能是在效率或稳健性方面,或者两者兼而有之。该项目旨在通过开发一类新颖的,可证明的和可扩展的半监督推理方法,为两个相当不同和活跃的研究领域的一系列基本问题提供答案:1)半监督设置中的因果推理,以及2)在标记中存在选择偏差的半监督推理。项目中概述的研究将在弥合现有文献中的一些主要差距方面取得进展,并提供对半监督推理及其微妙之处急需的统一理解。这些方法也将广泛适用于各个领域,例如精准医学和因果推理的生物医学研究。该项目还有一个重要的教育组成部分,包括研究生指导和通过短期课程开发课程,以提高对现代统计中这些令人兴奋的新领域的认识。在项目的第一部分,PI将在潜在结果框架下考虑半监督设置下的因果推理,并探索常见因果参数的半监督推理,例如平均治疗效果和分位数治疗效果,这两个参数在监督设置下得到了广泛的研究,但在半监督设置下很少得到研究。PI将致力于开发半监督方法,用于对这些参数进行所谓的双鲁棒估计,这些方法可以提高(如果不是最优)效率,以及比可实现的最佳监督对口物更强的鲁棒性。项目的第二部分将考虑半监督推理,其中标记机制具有固有的选择偏差,从而使标记和未标记的数据不均匀分布。这样的设置虽然具有很大的实际意义,但迄今为止很少得到解决,部分原因是它们的分析相当具有挑战性,因为标记分数衰减到零,导致自然违反所谓的正/重叠假设。在这种情况下,PI将通过双鲁棒估计方法探索对各种参数(例如平均响应和平均处理效果(在因果框架下))的有效和速率最优的半监督推理,以及用于估计衰减倾向得分的建模策略,这是一个不可避免的挑战,也是一个独立的兴趣。在整个过程中,PI的重点将是开发具有严格理论保证的方法,以及满足大型现代数据集上预期应用程序所需的可扩展性的有效实现。提出的方法还将汇集经典半参数推理和现代高维统计理论的工具和思想的协同作用。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project will develop methods for robust statistical inference in semi-supervised settings. Unlike more traditional data settings, semi-supervised settings are characterized by two types of available data: 1) a typical small or moderate sized labeled (or supervised) data containing observations for a response (or outcome) and a set of covariates (or predictors), and 2i) a much larger sized unlabeled (or unsupervised) data having observations only for the covariates. Such settings arise naturally whenever the covariates are easily available for a large cohort, while the response may be difficult and/or expensive to obtain due to practical constraints. These are increasingly relevant in modern studies in the big data era with large unlabeled databases (often electronically recorded) becoming easily available (and tractable) on top of a labeled data. Examples are ubiquitous across many disciplines, including computer science, machine learning, econometrics, and biomedical applications like electronic health records and integrative genomics. Statistical inference in semi-supervised settings is therefore of substantial interest. The ultimate question here is to investigate when and how one can use the extra information available from the large unlabeled data to “improve” upon a corresponding supervised approach, where improvement could be in terms of efficiency or robustness or both. This project aims to provide answers to such questions by developing a class of novel, provable and scalable semi-supervised inference methods for a range of fundamental problems in two fairly distinct and active research areas: 1) causal inference in semi-supervised settings, and 2) semi-supervised inference in the presence of selection bias in labeling. The research outlined in the project will lead to advances in bridging some major gaps in the existing literature and providing a much-needed unified understanding of semi-supervised inference and its subtleties. The methods will also have wide applicability to various domain areas, e.g. biomedical studies for precision medicine and causal inference. The project also has a significant education component, including mentoring of graduate students and curriculum development via short courses to raise awareness about these exciting new areas in modern statistics.In the first part of the project, the PI will consider causal inference in semi-supervised settings under the potential outcome framework, and explore semi-supervised inference for popular causal parameters, e.g. the average treatment effect and the quantile treatment effect, both of which have been widely studied in supervised settings but rarely so under semi-supervised settings. The PI will aim to develop semi-supervised methods for so-called doubly robust estimation of such parameters that can lead to improved (if not optimal) efficiency, as well as much stronger robustness properties than their best achievable supervised counterparts. The second part of the project will consider semi-supervised inference where the labeling mechanism has inherent selection bias, thus making the labeled and unlabeled data unequally distributed. Such settings, while of great practical relevance, have rarely been addressed so far, partly because their analysis is quite challenging since the labeling fraction decays to zero leading to a natural violation of the so-called positivity/overlap assumption. Under this setting, the PI will explore efficient and rate-optimal semi-supervised inference for various parameters, e.g. the mean response and the average treatment effect (under a causal framework), via doubly robust estimation methods, as well as modeling strategies for estimating the decaying propensity score which arises as an inevitable challenge and is of independent interest. Throughout, the PI's emphasis will be on developing methods with rigorous theoretical guarantees as well as efficient implementation that meets the scalability demanded by the intended applications on large modern datasets. The proposed methods will also bring together a synergy of tools and ideas from classical semi-parametric inference and modern high dimensional statistics theory.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Double robust semi-supervised inference for the mean: selection bias under MAR labeling with decaying overlap
均值的双重鲁棒半监督推理:具有衰减重叠的 MAR 标签下的选择偏差
DOI:
10.1093/imaiai/iaad021
发表时间:
2023
期刊:
Information and Inference: A Journal of the IMA
影响因子:
--
作者:
[Zhang, Yuqian, Chakrabortty, Abhishek, Bradic, Jelena]
通讯作者:
Bradic, Jelena
海外基金