Interpretation Attacks and Defenses on Predictive Models Using Electronic Health Records

Interpretation Attacks and Defenses on Predictive Models Using Electronic Health Records
复制标题

DOI:
10.1007/978-3-031-43418-1_27
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Fereshteh Razmi;Jian Lou;Yuan Hong;Li Xiong
Fereshteh Razmi;Jian Lou;Yuan Hong;Li Xiong
中科院分区:
其他
文献类型:
--
作者:
Fereshteh Razmi;Jian Lou;Yuan Hong;Li Xiong

文献摘要

相似文献

复杂的深度神经网络的出现使得采用解释方法来深入了解模型预测背后的原理变得至关重要。然而,最近的研究揭示了对这些解释的攻击,其目的是欺骗用户并破坏模型的可信度。它在医疗系统中尤其重要,因为解释在解释结果时至关重要。本文提出了第一个解释攻击预测模型使用顺序电子健康记录(EHR)。以前在图像解释方面的尝试主要利用基于梯度的方法,但我们的研究表明,我们的攻击可以在不依赖于模型梯度的EHR解释上取得重大成功。我们引入与EHR数据兼容的指标来评估攻击的成功。此外,我们的研究结果表明,成功识别传统对抗性示例的检测方法对我们的攻击无效。然后,我们提出了一种防御方法,利用自动编码器去噪的数据和提高解释的鲁棒性。我们的研究结果表明,这种去噪方法优于广泛使用的防御方法,SmoothGrad,这是基于添加噪声的数据。
The emergence of complex deep neural networks made it crucial to employ interpretation methods for gaining insight into the rationale behind model predictions. However, recent studies have revealed attacks on these interpretations, which aim to deceive users and subvert the trustworthiness of the models. It is especially critical in medical systems, where interpretations are essential in explaining outcomes. This paper presents the first interpretation attack on predictive models using sequential electronic health records (EHRs). Prior attempts in image interpretation mainly utilized gradient-based methods, yet our research shows that our attack can attain significant success on EHR interpretations that do not rely on model gradients. We introduce metrics compatible with EHR data to evaluate the attack’s success. Moreover, our findings demonstrate that detection methods that have successfully identified conventional adversarial examples are ineffective against our attack. We then propose a defense method utilizing auto-encoders to de-noise the data and improve the interpretations’ robustness. Our results indicate that this de-noising method outperforms the widely used defense method, SmoothGrad, which is based on adding noise to the data.