Collaborative Research: RI: Small: Post hoc Explanations in the Wild: Exposing Vulnerabilities and Ensuring Robustness
Collaborative Research: RI: Small: Post hoc Explanations in the Wild: Exposing Vulnerabilities and Ensuring Robustness
批准号:
2008956
负责人:
Sameer Singh
金额:
$22.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2023-09-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The successful adoption of machine learning (ML) models in critical domains such as healthcare and criminal justice relies heavily on how well decision makers are able to understand and trust the functionality of these models. However, the proprietary nature and increasing complexity of ML models makes it challenging for domain experts to understand these complex "black boxes". Consequently, there has been a recent surge in techniques that explain black box models in a human interpretable manner by approximating them using simpler models. However, it is unclear to what extent these post hoc explanation techniques may mislead end users by giving them a false sense of security, and luring them into trusting and deploying untrustworthy black boxes. This project will build rigorous frameworks to expose the vulnerabilities of existing explanation techniques, assess how these vulnerabilities can manifest in real world applications, and develop new techniques to defend against these vulnerabilities. This project has the potential to significantly speed up the adoption of ML in a variety of domains including criminal justice (e.g., bail decisions), health care (e.g., patient diagnosis and treatment), and financial lending (e.g., loan approval).The goal of this project is to characterize the vulnerabilities of existing explanation techniques, understand how adversaries can exploit these vulnerabilities, and develop techniques to defend against them. The project will focus on the following subtasks: 1) understanding the real-world consequences of misleading explanations by conducting user studies and detailed interviews with domain experts in healthcare and criminal justice 2) identifying critical vulnerabilities in state-of-the-art explanation techniques that can be exploited by adversarial entities to generate misleading explanations, and 3) developing novel techniques for building robust and reliable explanations that are not prone to these vulnerabilities and thereby provide domain experts and other stakeholders with faithful explanations of complex black box models. With these contributions, the project will initiate a new body of research in ML interpretability that focuses on understanding how adversaries can manipulate explanation techniques, and how to defend against such attacks.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2020-08
期刊:
影响因子:
--
作者:
[Dylan Slack;Sophie Hilgard;Sameer Singh;Himabindu Lakkaraju]
通讯作者:
Dylan Slack;Sophie Hilgard;Sameer Singh;Himabindu Lakkaraju
DOI:
10.18653/v1/2021.naacl-main.75
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
作者:
[Pouya Pezeshkpour;Sarthak Jain;Byron C. Wallace;Sameer Singh]
通讯作者:
Pouya Pezeshkpour;Sarthak Jain;Byron C. Wallace;Sameer Singh
DOI:
--
发表时间:
2021
期刊:
Advances in Neural Information Processing Systems (NeurIPS
影响因子:
--
作者:
[Slack, Dylan, Hilgard, Anna, Lakkaraju, Himabindu, Singh, Sameer]
通讯作者:
Singh, Sameer
DOI:
10.18653/v1/2022.findings-acl.153
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
作者:
[Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace]
通讯作者:
Pouya Pezeshkpour;Sarthak Jain;Sameer Singh;Byron C. Wallace
DOI:
--
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
作者:
[Himabindu Lakkaraju;Dylan Slack;Yuxin Chen;Chenhao Tan;Sameer Singh]
通讯作者:
Himabindu Lakkaraju;Dylan Slack;Yuxin Chen;Chenhao Tan;Sameer Singh
CAREER: Detecting, Understanding, and Fixing Vulnerabilities in Natural Language Processing Models
-
批准号:2046873
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2021
-
负责人:Sameer Singh
-
依托单位:
CCRI: ENS: Machine Learning Democratization via a Linked, Annotated Repository of Datasets
-
批准号:1925741
-
项目类别:Standard Grant
-
资助金额:$179.3万
-
财政年份:2019
-
负责人:Sameer Singh
-
依托单位:
CRII: RI: Explaining Decisions of Black-box Models via Input Perturbations
-
批准号:1756023
-
项目类别:Standard Grant
-
资助金额:$17.49万
-
财政年份:2018
-
负责人:Sameer Singh
-
依托单位:
RI: Small: Modeling Multiple Modalities for Knowledge-Base Construction
-
批准号:1817183
-
项目类别:Standard Grant
-
资助金额:$44.8万
-
财政年份:2018
-
负责人:Sameer Singh
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: