Precursory Analysis of Attack-Log Time Series by Machine Learning for Detecting Bots in CAPTCHA

Precursory Analysis of Attack-Log Time Series by Machine Learning for Detecting Bots in CAPTCHA
复制标题

DOI:
10.1109/icoin50884.2021.9333881
复制
发表时间:
2021-01
期刊:
2021 International Conference on Information Networking (ICOIN)
影响因子:
--
通讯作者:
Tsuyoshi Arai;Y. Okabe;Yoshinori Matsumoto
Tsuyoshi Arai;Y. Okabe;Yoshinori Matsumoto
中科院分区:
其他
文献类型:
--
作者:
Tsuyoshi Arai;Y. Okabe;Yoshinori Matsumoto

文献摘要

相似文献

Captcha(区分计算机和人类的全自动公共图灵测试)通常被用作避免机器人攻击网站的技术。最先进的验证码根据客户的行为在难度上有所不同,允许在不牺牲简单性的情况下有效地检测机器人。在本研究中,我们将重点放在利用有监督机器学习从过去的访问日志时间序列中检测僵尸程序。我们已经分析了几个Web服务的访问日志,这些Web服务使用基于云的商业CAPTCHA服务Capy Putle CAPTCHA。实验表明,仅通过对第一天的访问日志作为训练数据进行前兆分析,就可以在一个月以上的攻击中进行高精度的BOT检测。此外,我们对识别结果中发现的假阳性数据进行了人工分析,发现所提出的模型实际上检测到了机器人的访问,这是第一阶段人工识别标记准备训练数据时所忽略的。
CAPTCHA (Completely Automated Public Turing test to tell Computers and Humans Apart) is commonly utilized as a technology for avoiding attacks to Web sites by bots. State-of-the-art CAPTCHAs vary in difficulty based on the client’s behavior, allowing for efficient bot detection without sacrificing simplicity. In this research, we focus on detecting bots by supervised machine learning from access-log time series in the past. We have analysed access logs to several Web services which are using a commercial cloud-based CAPTCHA service, Capy Puzzle CAPTCHA. Experiments show that bot detection in attacks over a month can be performed with high accuracy by precursory analysis of the access log in only the first day as training data. In addition, we have manually analyzed the data that are found to be False Positive in the discrimination results, and it is found that the proposed model actually detects access by bots, which had been overlooked in the first-stage manual discrimination of flags in preparation of training data.