Platform-Oblivious Anti-Spam Gateway

Platform-Oblivious Anti-Spam Gateway
复制标题

DOI:
10.1145/3485832.3488024
复制
发表时间:
2021-12
期刊:
Proceedings of the 37th Annual Computer Security Applications Conference
影响因子:
--
通讯作者:
Yihe Zhang;Xu Yuan;N. Tzeng
Yihe Zhang;Xu Yuan;N. Tzeng
中科院分区:
其他
文献类型:
--
作者:
Yihe Zhang;Xu Yuan;N. Tzeng

文献摘要

被引文献

相似文献

本文提出了一种针对多个基于语言的社交平台的新型反垃圾邮件网关,以统一暴露其垃圾邮件消息的异常属性,以进行有效检测。我们没有标记真实数据集并提取关键特征,这些工作既费力又耗时,而是从目标数据中粗略地挖掘垃圾邮件和火腿的种子语料库(旨在垃圾邮件分类),然后将它们重建为参考。为了从语义和句法角度捕捉每个单词的丰富信息,我们利用自然语言处理(NLP)模型将每个单词嵌入到高维向量空间中,并使用神经网络来训练垃圾邮件单词模型。之后,使用该模型针对所有包含的词干词预测的垃圾邮件分数对每条消息进行编码。编码的消息由著名的异常值技术处理以产生各自的分数,使我们能够对它们进行排名以使异常值可见。我们的解决方案是无监督的,不依赖于任何平台或数据集的细节,是平台无关的。通过大量的实验,我们的解决方案被证明可以有效地暴露垃圾邮件发送者的异常特征,在几乎所有指标上都优于所有经过检查的无监督方法,甚至可能更好的监督方法。
This paper addresses a novel anti-spam gateway targeting multiple linguistic-based social platforms to expose the outlier property of their spam messages uniformly for effective detection. Instead of labeling ground truth datasets and extracting key features, which are labor-intensive and time-consuming, we start with coarsely mining seed corpora of spams and hams from the target data (aiming for spam classification), before reconstructing them as the reference. To catch each word’s rich information in the semantic and syntactic perspectives, we then leverage the natural language processing (NLP) model to embed each word into the high-dimensional vector space and use a neural network to train a spam word model. After that, each message is encoded by using the predicted spam scores from this model for all included stem words. The encoded messages are processed by the prominent outlier techniques to produce their respective scores, allowing us to rank them for making the outlier visible. Our solution is unsupervised, without relying on specifics of any platform or dataset, to be platform-oblivious. Through extensive experiments, our solution is demonstrated to expose spammers’ outlier characteristics effectively, outperform all examined unsupervised methods in almost all metrics, and may even better supervised counterparts.