Robust and Efficient Statistical Inference in Large Scale Semi-Supervised Settings
Robust and Efficient Statistical Inference in Large Scale Semi-Supervised Settings
批准号:
2113768
负责人:
Abhishek Chakrabortty
金额:
$17.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-01 至 2024-07-31
中文摘要
该项目将开发在半监督环境下进行稳健统计推断的方法。与更传统的数据设置不同,半监督设置的特征在于两种类型的可用数据:1)包含对响应(或结果)和一组协变量(或预测器)的观察的典型的小型或中等大小的标记(或监督)数据,以及2i)大得多的仅具有针对协变量的观察的大得多的未标记(或非监督)数据。每当协变量对于大的队列很容易获得时,这样的设置就会自然地出现,而响应可能由于实际限制而难以和/或昂贵地获得。这些在大数据时代的现代研究中变得越来越重要,大型的未标记数据库(通常是电子记录的)在已标记数据的基础上变得容易获得(和易于处理)。在许多学科中,例子随处可见,包括计算机科学、机器学习、计量经济学,以及电子健康记录和综合基因组学等生物医学应用程序。因此,在半监督环境中进行统计推断是非常有意义的。这里的最终问题是调查什么时候以及如何使用来自大量未标记数据的额外信息来对相应的监督方法进行“改进”,其中改进可能是在效率或健壮性方面,或者两者兼而有之。这个项目的目的是通过开发一类新颖的、可证明的和可扩展的半监督推理方法来解决这些问题,这些方法适用于两个相当不同和活跃的研究领域的基本问题:1)半监督环境下的因果推理;2)标签中存在选择偏差的半监督推理。该项目中概述的研究将导致在弥合现有文献中的一些主要差距方面取得进展,并提供对半监督推理及其微妙之处亟需的统一理解。这些方法还将广泛适用于各个领域,例如用于精确医学和因果推断的生物医学研究。该项目还有一个重要的教育部分,包括指导研究生和通过短期课程开发课程,以提高人们对现代统计学中这些令人兴奋的新领域的认识。在项目的第一部分,PI将在潜在结果框架下考虑半监督环境下的因果推理,并探索对常见因果参数的半监督推理,例如平均处理效果和分位数处理效果,这两个因素在监督环境中已被广泛研究,但在半监督环境下很少。PI的目标是开发半监督方法,用于这种参数的所谓双稳健估计,这些方法可以提高(如果不是最优的)效率,以及比它们最好的可实现的监督同行更强的稳健性。项目的第二部分将考虑半监督推理,其中标记机制具有固有的选择偏差,从而使已标记和未标记的数据不均匀分布。这种设置虽然具有很大的实际意义,但到目前为止很少被讨论,部分原因是他们的分析相当具有挑战性,因为标记分数衰减到零,导致自然违反所谓的积极/重叠假设。在这种情况下,PI将探索各种参数的有效和速率最优的半监督推理,例如平均响应和平均治疗效果(在因果框架下),通过双稳健估计方法以及估计衰减倾向分数的建模策略,这是一个不可避免的挑战和独立的兴趣。在整个过程中,PI的重点将是开发具有严格理论保证和有效实施的方法,以满足大型现代数据集上预期应用程序的可扩展性要求。建议的方法还将结合经典的半参数推理和现代高维统计理论的工具和思想。这一奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project will develop methods for robust statistical inference in semi-supervised settings. Unlike more traditional data settings, semi-supervised settings are characterized by two types of available data: 1) a typical small or moderate sized labeled (or supervised) data containing observations for a response (or outcome) and a set of covariates (or predictors), and 2i) a much larger sized unlabeled (or unsupervised) data having observations only for the covariates. Such settings arise naturally whenever the covariates are easily available for a large cohort, while the response may be difficult and/or expensive to obtain due to practical constraints. These are increasingly relevant in modern studies in the big data era with large unlabeled databases (often electronically recorded) becoming easily available (and tractable) on top of a labeled data. Examples are ubiquitous across many disciplines, including computer science, machine learning, econometrics, and biomedical applications like electronic health records and integrative genomics. Statistical inference in semi-supervised settings is therefore of substantial interest. The ultimate question here is to investigate when and how one can use the extra information available from the large unlabeled data to “improve” upon a corresponding supervised approach, where improvement could be in terms of efficiency or robustness or both. This project aims to provide answers to such questions by developing a class of novel, provable and scalable semi-supervised inference methods for a range of fundamental problems in two fairly distinct and active research areas: 1) causal inference in semi-supervised settings, and 2) semi-supervised inference in the presence of selection bias in labeling. The research outlined in the project will lead to advances in bridging some major gaps in the existing literature and providing a much-needed unified understanding of semi-supervised inference and its subtleties. The methods will also have wide applicability to various domain areas, e.g. biomedical studies for precision medicine and causal inference. The project also has a significant education component, including mentoring of graduate students and curriculum development via short courses to raise awareness about these exciting new areas in modern statistics.In the first part of the project, the PI will consider causal inference in semi-supervised settings under the potential outcome framework, and explore semi-supervised inference for popular causal parameters, e.g. the average treatment effect and the quantile treatment effect, both of which have been widely studied in supervised settings but rarely so under semi-supervised settings. The PI will aim to develop semi-supervised methods for so-called doubly robust estimation of such parameters that can lead to improved (if not optimal) efficiency, as well as much stronger robustness properties than their best achievable supervised counterparts. The second part of the project will consider semi-supervised inference where the labeling mechanism has inherent selection bias, thus making the labeled and unlabeled data unequally distributed. Such settings, while of great practical relevance, have rarely been addressed so far, partly because their analysis is quite challenging since the labeling fraction decays to zero leading to a natural violation of the so-called positivity/overlap assumption. Under this setting, the PI will explore efficient and rate-optimal semi-supervised inference for various parameters, e.g. the mean response and the average treatment effect (under a causal framework), via doubly robust estimation methods, as well as modeling strategies for estimating the decaying propensity score which arises as an inevitable challenge and is of independent interest. Throughout, the PI's emphasis will be on developing methods with rigorous theoretical guarantees as well as efficient implementation that meets the scalability demanded by the intended applications on large modern datasets. The proposed methods will also bring together a synergy of tools and ideas from classical semi-parametric inference and modern high dimensional statistics theory.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Double robust semi-supervised inference for the mean: selection bias under MAR labeling with decaying overlap
均值的双重鲁棒半监督推理:具有衰减重叠的 MAR 标签下的选择偏差
DOI:
10.1093/imaiai/iaad021
发表时间:
2023
期刊:
Information and Inference: A Journal of the IMA
影响因子:
--
作者:
[Zhang, Yuqian, Chakrabortty, Abhishek, Bradic, Jelena]
通讯作者:
Bradic, Jelena
海外基金