课题基金 / 基金详情

Extrapolative Analyses for Reliable Machine Learning Driven Scientific Discovery

Extrapolative Analyses for Reliable Machine Learning Driven Scientific Discovery
可靠的机器学习驱动的科学发现的外推分析
批准号:
2324394
负责人:
Junier Oliva
金额:
$59.46万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2026-08-31

项目摘要

项目成果

Junier Oliva的其他基金

相似基金

相关文献

中文摘要
翻译
在从实验和观测来源产生和分析数字数据方面出现了有希望的爆炸性增长,这为机器学习(ML)驱动的科学发现在诸如化学(化学信息学)和生物学(生物信息学)等高影响应用中提供了许多机会。不幸的是,当前的最大似然方法往往不能正确地描述与在训练中看到的明显不同的数据(即外推)。反过来,这又阻碍了我们做出真正超越我们现有知识的科学发现的能力。例如,这在化学虚拟筛选活动中具有重要意义,人们希望使用ML预测来指导昂贵的真实世界实验的潜在目标(例如,在药物发现应用中)。ML模型较差的外推能力可能会导致假阳性,浪费时间和资源,因为昂贵的合成和实验测试新的化学实体。这一奖项所产生的工作将提高ML模型在科学领域的实际效用,并防止错误使用模型预测。该项目还为研究生提供了研究培训机会。该项目开发了各种方法,以更准确地评估关于新输入的ML预测的可靠性,并提高模型的外推能力。首先,该项目开发了实证试验,以更准确地评估ML模型拟合过程在超出训练集分布支持的领域的外推能力。其次,利用外推评估,该项目开发了一些技术,以彻底探索可能外推的输入空间,以预测和过滤可能不可靠的预测。最后,该项目建立了指导获取新的训练数据的方法,一旦接受培训,将改进模型推断。该奖项由NSF高级网络基础设施办公室联合支持,由数学科学部颁发。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,认为值得支持。
英文摘要
There has been a promising explosion in the production and analysis of digital data from experimental and observational sources, which presents many opportunities for machine learning (ML) driven scientific discovery in high-impact applications such as chemistry (cheminformatics) and biology (bioinformatics). Unfortunately, current ML methodology often fails to properly characterize data markedly distinct from what was seen during training (i.e., extrapolation). This, in turn, hampers our ability to make scientific discoveries that truly extend past our current knowledge. For example, this is of great consequence in chemical virtual screening campaigns, where one hopes to use ML predictions to guide potential targets for expensive real-world experimentation (e.g., in drug discovery applications). Poor extrapolative power of ML models can result in false positives, wasting time and resources through costly synthesis and experimental testing of novel chemical entities. The work stemming from this award will improve the real-world utility of ML models in scientific domains and prevent the faulty use of model predictions. The project also provides research training opportunities for graduate students. This project develops various methodologies to more accurately assess the reliability of ML predictions on novel inputs and improve models' extrapolatory capabilities. First, the project develops empirical trials to more accurately evaluate the extrapolative capabilities of ML model fitting procedures on domains that lie beyond the training set distributional support. Second, utilizing extrapolative assessments, the project develops techniques to thoroughly explore the input space of possible extrapolation to anticipate and filter out likely unreliable predictions. Lastly, the project builds methodology to guide the acquisition of new training data that, once trained on, will improve model extrapolation.This award by the Division of Mathematical Sciences is jointly supported by the NSF Office of Advanced Cyberinfrastructure.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: From a Machine Detector to a Machine Detective: Decisions and Queries with Uncertain and Incomplete Information
海外基金