Dependable Predictive Inference with Uncertainty-Aware Machine Learning
Dependable Predictive Inference with Uncertainty-Aware Machine Learning
批准号:
2210637
负责人:
Matteo Sesia
金额:
$16.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-15 至 2025-07-31
中文摘要
复杂的统计和机器学习模型,包括深度神经网络,被广泛应用于许多领域,尽管人们对它们的可靠性感到严重担忧,但它们在数据驱动的科学中正变得越来越核心。这些模型并不总是可信的,特别是在敏感和高噪音的应用中,如基因组学中的应用,以及在所有机器学习预测将影响人们健康或福利的背景下。机器学习模型目前的一个重要局限性是,它们可能无法充分捕捉不确定性,而且它们的预测往往过于自信。此外,众所周知,机器学习模型有时会强化隐藏在数据中的潜在偏见,因此它们可能导致系统地对特定群体的个人存在偏见的预测。最后,许多统计和机器学习模型可能在训练它们的特定数据集中表现良好,但它们的预测对不断变化的数据环境并不稳健,例如对应于来自不同祖先的种群的个体的遗传分析的环境。为了解决上述限制,本研究项目将开发通用方法,以准确、公平和稳健地估计机器学习中的不确定性。在基因组学的特定背景下,这项工作将改进人类群体的遗传风险预测,促进个性化医学的进一步发展,弥合人群之间的健康差距,并有助于加深我们对可遗传疾病的科学知识。该项目将通过为研究生提供培训机会,支持统计和机器学习研究方面的教育。该项目还将通过帮助支持研究人员参与南加州大学的多样性、包容性和访问JumpStart倡议,促进统计和机器学习研究的多样性。特别是,调查员将为本科生提供暑期研究机会,重点关注本项目的主题。这项研究由三个不同但密切相关的部分组成。第一部分将开发新的保角推理方法来训练和校准既准确又可靠的不确定性感知机器学习模型。这项研究将涉及开发新的损失函数和创新的随机优化算法。该项目的第二部分将开发训练和校准不确定性感知机器学习模型的方法,这些模型公平地对待属于不同群体的个人,仔细地使用坚持观察来纠正可能的算法或数据偏差。该项目的第三部分将开发基于数据保持和保角推理的方法,以构建对协变量分布可能发生的变化更稳健的预测模型。这些模型将能够利用可用的预测变量之间可能的相互作用,并最终产生强大的可遗传疾病遗传风险的多变量模型,这些模型可能在不同的人群中依赖。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Complex statistical and machine learning models, including deep neural networks, are widely applied in many fields and they are becoming increasingly central to data-driven science, despite serious concerns about their reliability. These models cannot always be trusted, especially in sensitive and high-noise applications such as those found in genomics, as well as in all of those contexts in which machine learning predictions will affect people’s health or welfare. A crucial current limitation of machine learning models is that they may not adequately capture uncertainty and their predictions often tend to be overconfident. Further, machine learning models are known to sometimes reinforce latent biases hidden in the data, and thus they may lead to predictions that are systematically biased against certain groups of individuals. Finally, many statistical and machine learning models may perform well within the specific data set in which they are trained, but their predictions are not robust to changing data environments, such as those corresponding to the genetic analysis of individuals from populations with different ancestries. To address the above limitations, this research project will develop general methods for accurate, fair, and robust uncertainty estimation in machine learning. In the specific contexts of genomics, this work will lead to improved genetic risk prediction across human populations, facilitating further developments in personalized medicine, bridging health disparities across populations, and helping deepen our scientific knowledge of heritable diseases. This project will support education in statistical and machine learning research by providing training opportunities for graduate students. This project will also help promote diversity in statistical and machine learning research by helping support the investigator’s involvement with the Diversity, Inclusion, Access JumpStart initiative of the University of Southern California. In particular, the investigator will offer summer research opportunities focusing for undergraduate students on the topics of this project.This research consists of three distinct but closely connected parts. The first part will develop novel conformal inference methods to train and calibrate uncertainty-aware machine learning models that are both accurate and reliable. This research will involve the development of novel loss functions and innovative stochastic optimization algorithms. The second part of this project will develop methods for training and calibrating uncertainty-aware machine learning models that treat individuals belonging to different groups fairly, carefully using hold-out observations to correct for possible algorithmic or data biases. The third part of this project will develop methods based on data holdout and conformal inference to construct predictive models that are more robust to possible shifts in the covariate distribution. These models will be able to leverage possible interactions among the available predictive variables and ultimately lead to powerful multivariate models of genetic risk for heritable diseases that may be relied on across different populations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.48550/arxiv.2205.05878
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
作者:
[Bat-Sheva Einbinder;Yaniv Romano;Matteo Sesia;Yanfei Zhou]
通讯作者:
Bat-Sheva Einbinder;Yaniv Romano;Matteo Sesia;Yanfei Zhou
Conformal Frequency Estimation with Sketched Data
使用草图数据进行共形频率估计
DOI:
--
发表时间:
2022
期刊:
Advances in neural information processing systems
影响因子:
--
作者:
[Sesia, Matteo, Favaro, Stefano]
通讯作者:
Favaro, Stefano
海外基金