Evaluation of Federated Learning in Phishing Email Detection.

Evaluation of Federated Learning in Phishing Email Detection.
复制标题

DOI:
10.3390/s23094346
复制
发表时间:
2023-04-27
期刊:
Sensors (Basel, Switzerland)
影响因子:
--
通讯作者:
Zheng Y
Zheng Y
中科院分区:
其他
文献类型:
--
作者:
Thapa C;Tang JW;Abuadbba A;Gao Y;Camtepe S;Nepal S;Almashor M;Zheng Y

文献摘要

参考文献

被引文献

相似文献

使用人工智能(AI)来检测网络钓鱼电子邮件主要依赖于大规模的集中数据集,这导致了无数的隐私、信任和法律问题。此外,考虑到泄露商业敏感信息的风险,企业一直不愿共享电子邮件。因此,很难获得足够的电子邮件来有效地训练全球人工智能模型。因此,保护隐私的分布式和协作机器学习,特别是联邦学习(FL),是一个理想的选择。由于它在医疗保健领域已经很普遍,在多组织协作的背景下,基于fl的网络钓鱼检测的有效性和有效性仍然存在问题。据我们所知,这里的工作是第一次调查在网络钓鱼电子邮件检测中使用FL。本研究的重点是建立一个深度神经网络模型,特别是循环卷积神经网络(RNN)和来自变压器的双向编码器表示(BERT),用于网络钓鱼电子邮件检测。我们分析了不同设置下的fl纠缠学习性能,包括(i)组织之间平衡和不对称的数据分布以及(ii)可扩展性。我们的结果证实了FL在网络钓鱼邮件检测方面的性能统计数据与平衡数据集和低组织计数的集中式学习的可比性。此外,当增加组织数量时,我们观察到性能的变化。对于固定总数的电子邮件数据集,当组织数量从2个增加到10个时,基于rnn的全局模型的准确率下降了1.8%。相比之下,当组织计数从2增加到5时,BERT准确率提高了0.6%。然而,如果我们通过在FL框架中引入新的组织来增加整个电子邮件数据集,则通过实现更快的收敛速度来提高组织级别的性能。此外,如果电子邮件数据集分布高度不对称,由于输出高度不稳定,FL的整体全局模型性能会受到影响。
The use of artificial intelligence (AI) to detect phishing emails is primarily dependent on large-scale centralized datasets, which has opened it up to a myriad of privacy, trust, and legal issues. Moreover, organizations have been loath to share emails, given the risk of leaking commercially sensitive information. Consequently, it has been difficult to obtain sufficient emails to train a global AI model efficiently. Accordingly, privacy-preserving distributed and collaborative machine learning, particularly federated learning (FL), is a desideratum. As it is already prevalent in the healthcare sector, questions remain regarding the effectiveness and efficacy of FL-based phishing detection within the context of multi-organization collaborations. To the best of our knowledge, the work herein was the first to investigate the use of FL in phishing email detection. This study focused on building upon a deep neural network model, particularly recurrent convolutional neural network (RNN) and bidirectional encoder representations from transformers (BERT), for phishing email detection. We analyzed the FL-entangled learning performance in various settings, including (i) a balanced and asymmetrical data distribution among organizations and (ii) scalability. Our results corroborated the comparable performance statistics of FL in phishing email detection to centralized learning for balanced datasets and low organizational counts. Moreover, we observed a variation in performance when increasing the organizational counts. For a fixed total email dataset, the global RNN-based model had a 1.8% accuracy decrease when the organizational counts were increased from 2 to 10. In contrast, BERT accuracy increased by 0.6% when increasing organizational counts from 2 to 5. However, if we increased the overall email dataset by introducing new organizations in the FL framework, the organizational level performance improved by achieving a faster convergence speed. In addition, FL suffered in its overall global model performance due to highly unstable outputs if the email dataset distribution was highly asymmetric.
DOI: 10.1109/tdsc.2018.2864993
发表时间: 2018-11-01
影响因子: 7.3
作者:
Gutierrez, Christopher N.;Kim, Taegyu;Bagchi, Saurabh
通讯作者: Bagchi, Saurabh
DOI: 10.1109/comst.2019.2957750
发表时间: 2020-01-01
影响因子: 35.6
作者:
Das, Avisha;Baki, Shahryar;Dunbar, Arthur
通讯作者: Dunbar, Arthur
DOI: 10.1109/access.2019.2913705
发表时间: 2019-01-01
期刊: IEEE ACCESS
影响因子: 3.9
作者:
Fang, Yong;Zhang, Cheng;Yang, Yue
通讯作者: Yang, Yue
DOI: 10.1007/s11280-017-0524-3
发表时间: 2018-11-01
影响因子: 3.7
作者:
Gao, Jianliang;Ping, Qing;Wang, Jianxin
通讯作者: Wang, Jianxin
DOI: 10.3233/jcs-2010-0371
发表时间: 2010-01-01
影响因子: 1.2
作者:
Bergholz, Andre;De Beer, Jan;Strobel, Siehyun
通讯作者: Strobel, Siehyun