Learning from the Ones that Got Away: Detecting New Forms of Phishing Attacks
Learning from the Ones that Got Away: Detecting New Forms of Phishing Attacks
复制标题
DOI:
10.1109/tdsc.2018.2864993
复制
发表时间:
2018-11-01
影响因子:
7.3
通讯作者:
Bagchi, Saurabh
中科院分区:
文献类型:
--
作者:
Gutierrez, Christopher N.;Kim, Taegyu;Bagchi, Saurabh
Phishing attacks continue to pose a major threat for computer system defenders, often forming the first step in a multi-stage attack. There have been great strides made in phishing detection; however, some phishing emails appear to pass through filters by making simple structural and semantic changes to the messages. We tackle this problem through the use of a machine learning classifier operating on a large corpus of phishing and legitimate emails. We design SAF(E)-PC (Semi-Automated Feature generation for Phish Classification), a system to extract features, elevating some to higher level features, that are meant to defeat common phishing email detection strategies. To evaluate SAF(E)-PC, we collect a large corpus of phishing emails from the central IT organization at a tier-1 university. The execution of SAF(E)-PC on the dataset exposes hitherto unknown insights on phishing campaigns directed at university users. SAF(E)-PC detects more than 70 percent of the emails that had eluded our production deployment of Sophos, a state-of-the-art email filtering tool. It also outperforms SpamAssassin, a commonly used email filtering tool. We also developed an online version of SAF(E)-PC, that can be incrementally retrained with new samples. Its detection performance improves with time as new samples are collected, while the time to retrain the classifier stays constant.