Exploring Adversaries to Defend Audio CAPTCHA

Exploring Adversaries to Defend Audio CAPTCHA
复制标题

探索防御音频验证码的对手

DOI:
10.1109/icmla.2019.00192
复制
发表时间:
2019
期刊:
2019 18th IEEE International Conference On Machine Learning And Applications (ICMLA)
影响因子:
--
通讯作者:
Teng
Teng
中科院分区:
--
文献类型:
--
作者:
Heemany Shekhar;M. Moh;Teng

文献摘要

被引文献

相似文献

验证码是网站用来区分人类(有效用户)和机器人(攻击者)的一种基于Web的身份验证方法。Audio Captcha是一种可访问的验证码,旨在为色盲、盲人和近视用户等视觉残疾用户提供帮助。首先,利用机器学习(ML)和深度学习(DL)模型分析了当前音频验证码对攻击的安全性。每个音频验证码由五个、七个或十个随机数字[0-9]一个接一个地说出,在整个音频长度中伴随着不同的背景噪声。如果ML或DL模型能够正确识别所有语音数字并以正确的顺序出现在单个音频验证码中,我们认为该验证码已被破解,攻击成功。在整篇文章中,准确性指的是攻击模型成功破解音频验证码。攻击精度越高,音频验证码就越不安全。在我们的基线实验中,我们发现攻击模型可以破解没有背景噪声或中等背景噪声的音频验证码,具有任意数量的语音数字,准确率接近99%到100%。然而,背景噪声较高的音频验证码相对更安全,攻击准确率为85%。其次,我们提出可以利用对抗性例子算法的概念来创建一种新的音频验证码,该音频验证码对攻击具有更强的弹性。我们发现,即使在对新的对抗性音频数据进行重新训练后,攻击准确率仍然只有25%到36%。最后,我们探讨了不同的算法,如基本迭代方法(BIM)和深度愚人创建对抗性音频验证码的好处。我们发现,只要攻击者从每种敌意音频数据集中获得的样本少于45%,防御就能成功阻止攻击。
CAPTCHA is a web-based authentication method used by websites to distinguish between humans (valid users) and bots (attackers). Audio captcha is an accessible captcha meant for the visually disabled section of users such as color-blind, blind, near-sighted users. Firstly, this paper analyzes how secure current audio captchas are from attacks using machine learning (ML) and deep learning (DL) models. Each audio captcha is made up of five, seven or ten random digits[0-9] spoken one after the other along with varying background noise throughout the length of the audio. If the ML or DL model is able to correctly identify all spoken digits and in the correct order of occurance in a single audio captcha, we consider that captcha to be broken and the attack to be successful. Throughout the paper, accuracy refers to the attack model's success at breaking audio captchas. The higher the attack accuracy, the more unsecure the audio captchas are. In our baseline experiments, we found that attack models could break audio captchas that had no background noise or medium background noise with any number of spoken digits with nearly 99% to 100% accuracy. Whereas, audio captchas with high background noise were relatively more secure with attack accuracy of 85%. Secondly, we propose that the concepts of adversarial examples algorithms can be used to create a new kind of audio captcha that is more resilient towards attacks. We found that even after retraining the models on the new adversarial audio data, the attack accuracy remained as low as 25% to 36% only. Lastly, we explore the benefits of creating adversarial audio captcha through different algorithms such as Basic Iterative Method (BIM) and deepFool. We found that as long as the attacker has less than 45% sample from each kinds of adversarial audio datasets, the defense will be successful at preventing attacks.