Career: Building Models that Avoid Spurious Correlations through Interpretability and Representation Learning
Career: Building Models that Avoid Spurious Correlations through Interpretability and Representation Learning
批准号:
2145542
负责人:
Rajesh Ranganath
金额:
$54.67万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-01 至 2027-06-30
中文摘要
人工智能(AI)建模的进步使人工智能能够发现并使用各种信息来做出准确的预测。这些预测涉及到我们的日常生活,例如通过监测健康状况的健身可穿戴设备。有时,人工智能模型的预测使用了不稳定或虚假的信息。例如,当呈现在草地上的骆驼时,使用沙子来区分图像中是骆驼还是奶牛是不正确的。由于使用虚假信息而导致人工智能模型失败的例子存在于医疗保健等其他领域,在这些领域,人工智能模型可以根据数据的收集方式而不是数据中的生理信息进行预测。该项目旨在开发工具来帮助识别人工智能模型何时使用虚假信息,并开发工具来构建更好的人工智能模型,以避免使用虚假信息。这个项目的结果将是算法,适用于多种类型的数据和领域。该项目将通过人工智能在现实世界中的新讲座促进本科生和博士生的发展,并将通过项目开发的工具制作的人工智能模型的可视化来提高数据素养。在这个项目中有两个技术重点。第一个重点是人工智能模型的可解释性。该主旨旨在开发有助于识别虚假信息使用的方法,并且在给定输入中的虚假信息的知识的情况下,可以帮助识别对预测标签有用的输入中的语义信息。这种推力将适应学习解释的概念,它旨在训练一个函数来突出预测标签的输入的重要部分,以识别和减轻虚假信息的任务。第二部分构建了新的表征学习算法,用于构建避免使用虚假信息的模型。虚假信息包括在一系列数据生成分布中变化的变量之间的关系。这一重点将设法研究基于重加权的估计器和灵活模型的局限性,以及更强有力的假设的效用,例如残差的存在。它还将研究解决违反积极性所需的假设,以及如何进行表征学习,以避免多模态数据中的虚假信息。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Advances in artificial intelligence (AI) modeling have allowed AI to uncover and use all kinds of information to make accurate predictions. These predictions touch our day-to-day life, for example through fitness wearables that monitor health. Sometimes the predictions made by AI models make use of information that is unstable or spurious. For example, using sand to classify whether an image contains a camel versus a cow would be incorrect when presented with a camel in a grassy field. Examples of AI models failing because of the use of spurious information exist in other domains such as healthcare, where AI models can make predictions on the basis of how the data was collected rather than on the physiological information in the data. This project aims to develop tools to both help identify when AI models make use of spurious information and tools to build better AI models that avoid the use of spurious information. The results of this project will be algorithms that are applicable across several types of data and domains. The project will foster the development of undergraduate and PhD students through new lectures on AI for the real world and will promote data literacy through visualizations of AI models made by the tools the project will develop.There are two technical thrusts in this project. The first thrust focuses on interpretability of AI models. This thrust seeks to develop methods that can help identify the use of spurious information and, given knowledge of spurious information in an input, can help identify the semantic information in an input useful in predicting a label. This thrust will adapt the concept of learning to explain, which seeks to train a function to highlight the important part of an input for predicting a label, to the task of identifying and downweighing spurious information. The second thrust constructs new representation learning algorithms for building models that avoid the use of spurious information. Spurious information consists of relationships between variables that change across a family of data generating distributions. This thrust will seek to study the limits of reweighting-based estimators and flexible models and the utility of stronger assumptions, such as the existence of a residual. It will also study assumptions needed to address violations of positivity and how to do representation learning that avoids spurious information in multimodal data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2208.08579
发表时间:
2022-08
期刊:
Proceedings of machine learning research
影响因子:
--
作者:
[Mukund Sudarshan;A. Puli;Wesley Tansey;R. Ranganath]
通讯作者:
Mukund Sudarshan;A. Puli;Wesley Tansey;R. Ranganath
Where to Diffuse, How to Diffuse and How to Get Back: Automated Learning in Multivariate Diffusions
在哪里扩散,如何扩散以及如何返回:多元扩散中的自动学习
DOI:
--
发表时间:
2023
期刊:
International Conference on Learning Representations
影响因子:
--
作者:
[Singhal, Raghav, Goldstein, Mark, Ranganath, Rajesh]
通讯作者:
Ranganath, Rajesh
DOI:
10.48550/arxiv.2302.12893
发表时间:
2023-02
期刊:
影响因子:
--
作者:
[N. Jethani;A. Saporta;R. Ranganath]
通讯作者:
N. Jethani;A. Saporta;R. Ranganath
DOI:
10.48550/arxiv.2302.04132
发表时间:
2023-02
期刊:
Proceedings of the ... AAAI Conference on Artificial Intelligence. AAAI Conference on Artificial Intelligence
影响因子:
--
作者:
[Lily H. Zhang;R. Ranganath]
通讯作者:
Lily H. Zhang;R. Ranganath
DOI:
10.48550/arxiv.2208.10759
发表时间:
2022-08
期刊:
Proceedings of machine learning research
影响因子:
--
作者:
[Xintian Han;Mark Goldstein;R. Ranganath]
通讯作者:
Xintian Han;Mark Goldstein;R. Ranganath
国内基金
海外基金
基于支链淀粉building blocks构建优质BE突变酶定向修饰淀粉调控机制的研究
-
批准号:31771933
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:郭丽
-
依托单位: