I am Robot: (Deep) Learning to Break Semantic Image CAPTCHAs

I am Robot: (Deep) Learning to Break Semantic Image CAPTCHAs
复制标题

DOI:
10.1109/eurosp.2016.37
复制
发表时间:
2016-03
期刊:
2016 IEEE European Symposium on Security and Privacy (EuroS&P)
影响因子:
--
通讯作者:
Suphannee Sivakorn;Iasonas Polakis;A. Keromytis
Suphannee Sivakorn;Iasonas Polakis;A. Keromytis
中科院分区:
其他
文献类型:
--
作者:
Suphannee Sivakorn;Iasonas Polakis;A. Keromytis

文献摘要

被引文献

相似文献

自成立以来,验证码已被广泛用于防止欺诈者进行非法活动。然而,经济激励导致了一场军备竞赛,欺诈者开发了自动化解决方案,反过来,验证码服务调整了他们的设计来破坏解决方案。然而,最近的工作提出了一种通用攻击,可以应用于任何基于文本的验证码方案。谷歌最近发布了最新版本的reCaptcha。他们的新系统的目标是双重的,最大限度地减少合法用户的工作,同时要求比文本识别对计算机更具挑战性的任务。ReCaptcha由“高级风险分析系统”驱动,该系统评估请求并选择将返回的验证码的难度。用户可能需要点击复选框,或通过识别具有类似内容的图像来解决挑战。在本文中,我们对reCaptcha进行了全面的研究,并探讨了请求的各个方面如何影响风险分析过程。通过大量的实验,我们发现了一些缺陷,这些缺陷使对手能够毫不费力地影响风险分析,绕过限制,并部署大规模攻击。随后,我们设计了一种新的低成本攻击,利用深度学习技术对图像进行语义注释。我们的系统非常有效,自动解决了70.78%的图像reCaptcha挑战,而每个挑战只需要19秒。我们还将我们的攻击应用于Facebook图像验证码,并实现了83.5%的准确率。根据我们的实验结果,我们提出了一系列的保护措施和修改,以影响我们的攻击的可扩展性和准确性。总的来说,虽然我们的研究重点是reCaptcha,但我们的发现具有广泛的影响,因为通过图像传达的语义信息越来越多地在自动推理领域内,captchas的未来依赖于对新方向的探索。
Since their inception, captchas have been widely used for preventing fraudsters from performing illicit actions. Nevertheless, economic incentives have resulted in an arms race, where fraudsters develop automated solvers and, in turn, captcha services tweak their design to break the solvers. Recent work, however, presented a generic attack that can be applied to any text-based captcha scheme. Fittingly, Google recently unveiled the latest version of reCaptcha. The goal of their new system is twofold, to minimize the effort for legitimate users, while requiring tasks that are more challenging to computers than text recognition. ReCaptcha is driven by an "advanced risk analysis system" that evaluates requests and selects the difficulty of the captcha that will be returned. Users may be required to click in a checkbox, or solve a challenge by identifying images with similar content. In this paper, we conduct a comprehensive study of reCaptcha, and explore how the risk analysis process is influenced by each aspect of the request. Through extensive experimentation, we identify flaws that allow adversaries to effortlessly influence the risk analysis, bypass restrictions, and deploy large-scale attacks. Subsequently, we design a novel low-cost attack that leverages deep learning technologies for the semantic annotation of images. Our system is extremely effective, automatically solving 70.78% of the image reCaptcha challenges, while requiring only 19 seconds per challenge. We also apply our attack to the Facebook image captcha and achieve an accuracy of 83.5%. Based on our experimental findings, we propose a series of safeguards and modifications for impacting the scalability and accuracy of our attacks. Overall, while our study focuses on reCaptcha, our findings have wide implications, as the semantic information conveyed via images is increasingly within the realm of automated reasoning, the future of captchas relies on the exploration of novel directions.