Dependable Predictive Inference with Uncertainty-Aware Machine Learning
Dependable Predictive Inference with Uncertainty-Aware Machine Learning
批准号:
2210637
负责人:
Matteo Sesia
金额:
$16.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-15 至 2025-07-31
中文摘要
复杂的统计和机器学习模型,包括深度神经网络,广泛应用于许多领域,它们正成为数据驱动科学的核心,尽管它们的可靠性受到严重关注。这些模型并不总是可信的,特别是在敏感和高噪声的应用中,例如基因组学中的应用,以及机器学习预测将影响人们健康或福利的所有背景下。机器学习模型目前的一个关键限制是,它们可能无法充分捕捉不确定性,而且它们的预测往往过于自信。此外,已知机器学习模型有时会强化隐藏在数据中的潜在偏见,因此它们可能导致对某些群体有系统偏见的预测。最后,许多统计和机器学习模型可能在训练它们的特定数据集内表现良好,但它们的预测对于不断变化的数据环境并不稳健,例如与来自不同祖先群体的个体的遗传分析相对应的预测。为了解决上述限制,本研究项目将开发机器学习中准确,公平和鲁棒的不确定性估计的通用方法。在基因组学的特定背景下,这项工作将改善对人类群体的遗传风险预测,促进个性化医疗的进一步发展,弥合人群之间的健康差距,并帮助加深我们对遗传性疾病的科学知识。该项目将通过为研究生提供培训机会,支持统计和机器学习研究方面的教育。该项目还将通过帮助支持调查人员参与南加州大学的多样性,包容性,访问快速启动计划,帮助促进统计和机器学习研究的多样性。特别是,研究者将为本科生提供暑期研究机会,重点关注本项目的主题。本研究由三个不同但密切相关的部分组成。第一部分将开发新的共形推理方法来训练和校准准确可靠的不确定性感知机器学习模型。这项研究将涉及新的损失函数和创新的随机优化算法的发展。该项目的第二部分将开发用于训练和校准不确定性感知机器学习模型的方法,这些模型公平地对待属于不同群体的个体,仔细使用保留观察来纠正可能的算法或数据偏差。该项目的第三部分将开发基于数据保持和共形推理的方法,以构建对协变量分布中可能的变化更鲁棒的预测模型。这些模型将能够利用可用的预测变量之间可能的相互作用,并最终导致强大的多变量模型的遗传风险的遗传性疾病,可能依赖于在不同的population.This奖项反映了NSF的法定使命,并已被认为是值得的支持,通过评估使用基金会的智力价值和更广泛的影响审查标准。
英文摘要
Complex statistical and machine learning models, including deep neural networks, are widely applied in many fields and they are becoming increasingly central to data-driven science, despite serious concerns about their reliability. These models cannot always be trusted, especially in sensitive and high-noise applications such as those found in genomics, as well as in all of those contexts in which machine learning predictions will affect people’s health or welfare. A crucial current limitation of machine learning models is that they may not adequately capture uncertainty and their predictions often tend to be overconfident. Further, machine learning models are known to sometimes reinforce latent biases hidden in the data, and thus they may lead to predictions that are systematically biased against certain groups of individuals. Finally, many statistical and machine learning models may perform well within the specific data set in which they are trained, but their predictions are not robust to changing data environments, such as those corresponding to the genetic analysis of individuals from populations with different ancestries. To address the above limitations, this research project will develop general methods for accurate, fair, and robust uncertainty estimation in machine learning. In the specific contexts of genomics, this work will lead to improved genetic risk prediction across human populations, facilitating further developments in personalized medicine, bridging health disparities across populations, and helping deepen our scientific knowledge of heritable diseases. This project will support education in statistical and machine learning research by providing training opportunities for graduate students. This project will also help promote diversity in statistical and machine learning research by helping support the investigator’s involvement with the Diversity, Inclusion, Access JumpStart initiative of the University of Southern California. In particular, the investigator will offer summer research opportunities focusing for undergraduate students on the topics of this project.This research consists of three distinct but closely connected parts. The first part will develop novel conformal inference methods to train and calibrate uncertainty-aware machine learning models that are both accurate and reliable. This research will involve the development of novel loss functions and innovative stochastic optimization algorithms. The second part of this project will develop methods for training and calibrating uncertainty-aware machine learning models that treat individuals belonging to different groups fairly, carefully using hold-out observations to correct for possible algorithmic or data biases. The third part of this project will develop methods based on data holdout and conformal inference to construct predictive models that are more robust to possible shifts in the covariate distribution. These models will be able to leverage possible interactions among the available predictive variables and ultimately lead to powerful multivariate models of genetic risk for heritable diseases that may be relied on across different populations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.48550/arxiv.2205.05878
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
作者:
[Bat-Sheva Einbinder;Yaniv Romano;Matteo Sesia;Yanfei Zhou]
通讯作者:
Bat-Sheva Einbinder;Yaniv Romano;Matteo Sesia;Yanfei Zhou
Conformal Frequency Estimation with Sketched Data
使用草图数据进行共形频率估计
DOI:
--
发表时间:
2022
期刊:
Advances in neural information processing systems
影响因子:
--
作者:
[Sesia, Matteo, Favaro, Stefano]
通讯作者:
Favaro, Stefano
海外基金