A Generative Adversarial Learning Framework for Breaking Text-Based CAPTCHA in the Dark Web

A Generative Adversarial Learning Framework for Breaking Text-Based CAPTCHA in the Dark Web
复制标题

DOI:
10.1109/isi49825.2020.9280537
复制
发表时间:
2020-11
期刊:
2020 IEEE International Conference on Intelligence and Security Informatics (ISI)
影响因子:
--
通讯作者:
Ning Zhang;Mohammadreza Ebrahimi;Weifeng Li;Hsinchun Chen
Ning Zhang;Mohammadreza Ebrahimi;Weifeng Li;Hsinchun Chen
中科院分区:
其他
文献类型:
--
作者:
Ning Zhang;Mohammadreza Ebrahimi;Weifeng Li;Hsinchun Chen

文献摘要

相似文献

网络威胁情报(CTI)需要对黑暗网络平台(例如,黑暗网络市场和纸牌商店)进行大规模的自动监控。虽然已有从表面网收集数据的现有方法,但大规模的暗网数据收集通常受到反爬行措施的阻碍。基于文本的验证码是这些措施中最令人望而却步的一种。基于文本的验证码要求用户识别难以阅读的字符组合。暗网验证码图案被故意设计为具有额外的背景噪音和可变的字符长度,以防止自动验证码中断。现有的验证码破解方法无法解决这些挑战,因此不适用于暗网。在这项研究中,我们提出了一种在黑暗网络中破解基于文本的验证码的新框架。该框架利用生成性对抗性网络(GAN)来抵消特定于暗网页的背景噪声,并利用一种增强的字符分割算法。我们提出的方法在Benchmark和Dark Web CAPTCHA试验台上进行了评估。该方法在所有数据集上的性能明显优于最新的基线方法,在暗网测试床上的成功率超过92.08%。我们的研究使CTI社区能够开发大规模暗网监控的高级能力。
Cyber threat intelligence (CTI) necessitates automated monitoring of dark web platforms (e.g., Dark Net Markets and carding shops) on a large scale. While there are existing methods for collecting data from the surface web, large-scale dark web data collection is commonly hindered by anti-crawling measures. Text-based CAPTCHA serves as the most prohibitive type of these measures. Text-based CAPTCHA requires the user to recognize a combination of hard-to-read characters. Dark web CAPTCHA patterns are intentionally designed to have additional background noise and variable character length to prevent automated CAPTCHA breaking. Existing CAPTCHA breaking methods cannot remedy these challenges and are therefore not applicable to the dark web. In this study, we propose a novel framework for breaking text-based CAPTCHA in the dark web. The proposed framework utilizes Generative Adversarial Network (GAN) to counteract dark web-specific background noise and leverages an enhanced character segmentation algorithm. Our proposed method was evaluated on both benchmark and dark web CAPTCHA testbeds. The proposed method significantly outperformed the state-of-the-art baseline methods on all datasets, achieving over 92.08% success rate on dark web testbeds. Our research enables the CTI community to develop advanced capabilities of large-scale dark web monitoring.