课题基金 / 基金详情

Collaborative Research: RI: Small: Post hoc Explanations in the Wild: Exposing Vulnerabilities and Ensuring Robustness

Collaborative Research: RI: Small: Post hoc Explanations in the Wild: Exposing Vulnerabilities and Ensuring Robustness
合作研究:RI:小型:事后解释:暴露漏洞并确保稳健性
批准号:
2008461
负责人:
Himabindu Lakkaraju
金额:
$22.5万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-10-01 至 2024-09-30

项目摘要

项目成果

Himabindu Lakkaraju的其他基金

相似基金

相关文献

中文摘要
翻译
在医疗保健和刑事司法等关键领域成功采用机器学习(ML)模型在很大程度上取决于决策者能够理解和信任这些模型的功能。然而,ML模型的专有性质和日益增加的复杂性使得领域专家难以理解这些复杂的“黑匣子”。因此,最近出现了一种技术,通过使用更简单的模型来近似它们,以人类可解释的方式解释黑盒模型。然而,目前还不清楚这些事后解释技术可能会在多大程度上误导最终用户,给他们一种虚假的安全感,并诱使他们信任和部署不可信的黑盒。该项目将建立严格的框架来暴露现有解释技术的漏洞,评估这些漏洞如何在真实的世界应用程序中表现出来,并开发新的技术来防御这些漏洞。该项目有可能大大加快ML在包括刑事司法在内的各个领域的采用(例如,保释决定),医疗保健(例如,患者诊断和治疗),以及金融贷款(例如,该项目的目标是描述现有解释技术的漏洞,了解对手如何利用这些漏洞,并开发防御这些漏洞的技术。该项目将侧重于以下子任务:1)通过进行用户研究和与医疗保健和刑事司法领域专家的详细访谈,了解误导性解释的现实后果2)识别最先进解释技术中的关键漏洞,这些漏洞可以被敌对实体利用来生成误导性解释,以及3)开发用于构建不容易出现这些漏洞的鲁棒且可靠的解释的新颖技术,从而为领域专家和其他利益相关者提供复杂黑盒模型的忠实解释。有了这些贡献,该项目将启动一个新的机器学习可解释性研究机构,重点是了解对手如何操纵解释技术,以及如何抵御此类攻击。该奖项反映了NSF的法定使命,并被认为值得通过使用基金会的智力价值和更广泛的影响审查标准进行评估来支持。
英文摘要
The successful adoption of machine learning (ML) models in critical domains such as healthcare and criminal justice relies heavily on how well decision makers are able to understand and trust the functionality of these models. However, the proprietary nature and increasing complexity of ML models makes it challenging for domain experts to understand these complex "black boxes". Consequently, there has been a recent surge in techniques that explain black box models in a human interpretable manner by approximating them using simpler models. However, it is unclear to what extent these post hoc explanation techniques may mislead end users by giving them a false sense of security, and luring them into trusting and deploying untrustworthy black boxes. This project will build rigorous frameworks to expose the vulnerabilities of existing explanation techniques, assess how these vulnerabilities can manifest in real world applications, and develop new techniques to defend against these vulnerabilities. This project has the potential to significantly speed up the adoption of ML in a variety of domains including criminal justice (e.g., bail decisions), health care (e.g., patient diagnosis and treatment), and financial lending (e.g., loan approval).The goal of this project is to characterize the vulnerabilities of existing explanation techniques, understand how adversaries can exploit these vulnerabilities, and develop techniques to defend against them. The project will focus on the following subtasks: 1) understanding the real-world consequences of misleading explanations by conducting user studies and detailed interviews with domain experts in healthcare and criminal justice 2) identifying critical vulnerabilities in state-of-the-art explanation techniques that can be exploited by adversarial entities to generate misleading explanations, and 3) developing novel techniques for building robust and reliable explanations that are not prone to these vulnerabilities and thereby provide domain experts and other stakeholders with faithful explanations of complex black box models. With these contributions, the project will initiate a new body of research in ML interpretability that focuses on understanding how adversaries can manipulate explanation techniques, and how to defend against such attacks.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(18)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3514094.3534159
发表时间: 2022-05
期刊: Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society
影响因子: --
作者: [Jessica Dai;Sohini Upadhyay;U. Aïvodji;Stephen H. Bach;Himabindu Lakkaraju]
通讯作者: Jessica Dai;Sohini Upadhyay;U. Aïvodji;Stephen H. Bach;Himabindu Lakkaraju
DOI: --
发表时间: 2021-06
期刊:
影响因子: --
作者: [Martin Pawelczyk;Chirag Agarwal;Shalmali Joshi;Sohini Upadhyay;Himabindu Lakkaraju]
通讯作者: Martin Pawelczyk;Chirag Agarwal;Shalmali Joshi;Sohini Upadhyay;Himabindu Lakkaraju
DOI: --
发表时间: 2020-08
期刊:
影响因子: --
作者: [Dylan Slack;Sophie Hilgard;Sameer Singh;Himabindu Lakkaraju]
通讯作者: Dylan Slack;Sophie Hilgard;Sameer Singh;Himabindu Lakkaraju
Towards Robust and Reliable Recourse
迈向稳健可靠的追索权
DOI: --
发表时间: 2021
期刊: Advances in neural information processing systems
影响因子: --
作者: [Upadhyay, Sohini, Joshi, Shalmali, Lakkaraju, Himabindu]
通讯作者: Lakkaraju, Himabindu
共 15 条
    Career: Towards a Systematic Characterization of Model Explanations for High-Stakes Decision Making
    • 批准号:
      2238714
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $55.07万
    • 财政年份:
      2023
    • 负责人:
      Himabindu Lakkaraju
    • 依托单位:
    国内基金
    海外基金
    Research on Quantum Field Theory without a Lagrangian Description
    • 批准号:
      24ZR1403900
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
      SATOSHI NAWATA
    • 依托单位:
    Cell Research
    Cell Research
    Cell Research (细胞研究)