Career: Building Models that Avoid Spurious Correlations through Interpretability and Representation Learning
Career: Building Models that Avoid Spurious Correlations through Interpretability and Representation Learning
批准号:
2145542
负责人:
Rajesh Ranganath
金额:
$54.67万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2027-06-30
中文摘要
人工智能(AI)建模的进步使AI能够发现并使用各种信息来做出准确的预测。这些预测影响了我们的日常生活,例如,通过监测健康的健身可穿戴设备。有时,人工智能模型做出的预测使用了不稳定或虚假的信息。例如,当在草地上展示骆驼时,用沙子来区分图像中是否包含骆驼和奶牛是不正确的。在医疗等其他领域,人工智能模型因使用虚假信息而失败的例子也存在,在这些领域,人工智能模型可以根据数据收集的方式而不是数据中的生理信息做出预测。该项目旨在开发工具,以帮助识别人工智能模型何时利用虚假信息,并开发工具来构建更好的人工智能模型,以避免使用虚假信息。该项目的结果将是适用于多种类型的数据和域的算法。该项目将通过针对现实世界的人工智能的新讲座来促进本科生和博士生的发展,并将通过该项目将开发的工具制作的人工智能模型的可视化来促进数据素养。该项目有两个技术要点。第一个重点是人工智能模型的可解释性。这一努力旨在开发能够帮助识别虚假信息的使用的方法,并且在给定输入中虚假信息的知识的情况下,可以帮助识别输入中对预测标签有用的语义信息。这一推力将使学习解释的概念适应于识别和降低虚假信息的任务,该概念寻求训练一种功能,以突出用于预测标签的输入的重要部分。第二个推力构建了新的表示学习算法,用于建立避免使用虚假信息的模型。虚假信息由变量之间的关系组成,这些变量在一系列数据生成分布中发生变化。这一努力将寻求研究基于重新加权的估计器和灵活模型的局限性,以及更强有力的假设的实用性,例如残差的存在。它还将研究解决正性违规所需的假设,以及如何进行表征学习,以避免多模式数据中的虚假信息。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Advances in artificial intelligence (AI) modeling have allowed AI to uncover and use all kinds of information to make accurate predictions. These predictions touch our day-to-day life, for example through fitness wearables that monitor health. Sometimes the predictions made by AI models make use of information that is unstable or spurious. For example, using sand to classify whether an image contains a camel versus a cow would be incorrect when presented with a camel in a grassy field. Examples of AI models failing because of the use of spurious information exist in other domains such as healthcare, where AI models can make predictions on the basis of how the data was collected rather than on the physiological information in the data. This project aims to develop tools to both help identify when AI models make use of spurious information and tools to build better AI models that avoid the use of spurious information. The results of this project will be algorithms that are applicable across several types of data and domains. The project will foster the development of undergraduate and PhD students through new lectures on AI for the real world and will promote data literacy through visualizations of AI models made by the tools the project will develop.There are two technical thrusts in this project. The first thrust focuses on interpretability of AI models. This thrust seeks to develop methods that can help identify the use of spurious information and, given knowledge of spurious information in an input, can help identify the semantic information in an input useful in predicting a label. This thrust will adapt the concept of learning to explain, which seeks to train a function to highlight the important part of an input for predicting a label, to the task of identifying and downweighing spurious information. The second thrust constructs new representation learning algorithms for building models that avoid the use of spurious information. Spurious information consists of relationships between variables that change across a family of data generating distributions. This thrust will seek to study the limits of reweighting-based estimators and flexible models and the utility of stronger assumptions, such as the existence of a residual. It will also study assumptions needed to address violations of positivity and how to do representation learning that avoids spurious information in multimodal data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2208.08579
发表时间:
2022-08
期刊:
Proceedings of machine learning research
影响因子:
--
作者:
[Mukund Sudarshan;A. Puli;Wesley Tansey;R. Ranganath]
通讯作者:
Mukund Sudarshan;A. Puli;Wesley Tansey;R. Ranganath
Where to Diffuse, How to Diffuse and How to Get Back: Automated Learning in Multivariate Diffusions
在哪里扩散,如何扩散以及如何返回:多元扩散中的自动学习
DOI:
--
发表时间:
2023
期刊:
International Conference on Learning Representations
影响因子:
--
作者:
[Singhal, Raghav, Goldstein, Mark, Ranganath, Rajesh]
通讯作者:
Ranganath, Rajesh
DOI:
10.48550/arxiv.2302.12893
发表时间:
2023-02
期刊:
影响因子:
--
作者:
[N. Jethani;A. Saporta;R. Ranganath]
通讯作者:
N. Jethani;A. Saporta;R. Ranganath
DOI:
10.48550/arxiv.2302.04132
发表时间:
2023-02
期刊:
Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence
影响因子:
--
作者:
[Lily H. Zhang;R. Ranganath]
通讯作者:
Lily H. Zhang;R. Ranganath
DOI:
10.48550/arxiv.2208.10759
发表时间:
2022-08
期刊:
Proceedings of machine learning research
影响因子:
--
作者:
[Xintian Han;Mark Goldstein;R. Ranganath]
通讯作者:
Xintian Han;Mark Goldstein;R. Ranganath
国内基金
海外基金
基于支链淀粉building blocks构建优质BE突变酶定向修饰淀粉调控机制的研究
-
批准号:31771933
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:郭丽
-
依托单位: