Counteracting Dark Web Text-Based CAPTCHA with Generative Adversarial Learning for Proactive Cyber Threat Intelligence

Counteracting Dark Web Text-Based CAPTCHA with Generative Adversarial Learning for Proactive Cyber Threat Intelligence
复制标题

DOI:
10.1145/3505226
复制
发表时间:
2022-01
期刊:
ACM Transactions on Management Information Systems (TMIS)
影响因子:
--
通讯作者:
Ning Zhang;Mohammadreza Ebrahimi;Weifeng Li;Hsinchun Chen
Ning Zhang;Mohammadreza Ebrahimi;Weifeng Li;Hsinchun Chen
中科院分区:
其他
文献类型:
--
作者:
Ning Zhang;Mohammadreza Ebrahimi;Weifeng Li;Hsinchun Chen

文献摘要

相似文献

大规模自动监控暗网(DW)平台是开发主动网络威胁情报(CTI)的第一步。虽然有从表面网络收集数据的有效方法,但大规模的暗网数据收集往往受到反抓取措施的阻碍。特别是,基于文本的CAPTCHA是暗网中最普遍和禁止的措施。基于文本的CAPTCHA通过强制用户输入难以识别的字母数字字符的组合来识别和阻止自动爬虫。在暗网中,CAPTCHA图像经过精心设计,带有额外的背景噪音和可变字符长度,以防止自动破坏CAPTCHA。现有的自动化CAPTCHA破解方法在克服这些暗网挑战方面存在困难。因此,解决基于暗网文本的CAPTCHA在很大程度上依赖于人工参与,这是劳动密集型和耗时的。在这项研究中,我们提出了一个新的框架,用于自动破解暗网CAPTCHA,以促进暗网数据收集。该框架包含一种新的生成方法,用于识别具有噪声背景和可变字符长度的基于暗网文本的CAPTCHA。为了消除对人类参与的需求,该框架利用生成对抗网络(GAN)来抵消暗网背景噪声,并利用增强的字符分割算法来处理具有可变字符长度的CAPTCHA图像。我们提出的框架DW-GAN在多个暗网CAPTCHA测试平台上进行了系统评估。DW-GAN在所有数据集上的表现都明显优于最先进的基准方法,在仔细收集的真实世界暗网数据集上实现了超过94.4%的成功率。我们进一步对一个新兴的暗网市场(DNM)进行了案例研究,以证明DW-GAN通过不超过三次的尝试自动解决CAPTCHA挑战来消除人类参与。我们的研究使CTI社区能够开发先进的大规模暗网监控。我们将DW-GAN代码作为GitHub中的开源工具提供给社区。
Automated monitoring of dark web (DW) platforms on a large scale is the first step toward developing proactive Cyber Threat Intelligence (CTI). While there are efficient methods for collecting data from the surface web, large-scale dark web data collection is often hindered by anti-crawling measures. In particular, text-based CAPTCHA serves as the most prevalent and prohibiting type of these measures in the dark web. Text-based CAPTCHA identifies and blocks automated crawlers by forcing the user to enter a combination of hard-to-recognize alphanumeric characters. In the dark web, CAPTCHA images are meticulously designed with additional background noise and variable character length to prevent automated CAPTCHA breaking. Existing automated CAPTCHA breaking methods have difficulties in overcoming these dark web challenges. As such, solving dark web text-based CAPTCHA has been relying heavily on human involvement, which is labor-intensive and time-consuming. In this study, we propose a novel framework for automated breaking of dark web CAPTCHA to facilitate dark web data collection. This framework encompasses a novel generative method to recognize dark web text-based CAPTCHA with noisy background and variable character length. To eliminate the need for human involvement, the proposed framework utilizes Generative Adversarial Network (GAN) to counteract dark web background noise and leverages an enhanced character segmentation algorithm to handle CAPTCHA images with variable character length. Our proposed framework, DW-GAN, was systematically evaluated on multiple dark web CAPTCHA testbeds. DW-GAN significantly outperformed the state-of-the-art benchmark methods on all datasets, achieving over 94.4% success rate on a carefully collected real-world dark web dataset. We further conducted a case study on an emergent Dark Net Marketplace (DNM) to demonstrate that DW-GAN eliminated human involvement by automatically solving CAPTCHA challenges with no more than three attempts. Our research enables the CTI community to develop advanced, large-scale dark web monitoring. We make DW-GAN code available to the community as an open-source tool in GitHub.