CAREER: Machine Learning Assisted Crowdsourcing for Phishing Defense
CAREER: Machine Learning Assisted Crowdsourcing for Phishing Defense
批准号:
1750101
负责人:
Gang Wang
金额:
$53.85万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-01 至 2020-06-30
中文摘要
该项目旨在通过结合人类和机器智能来应对日益增长的钓鱼攻击威胁。钓鱼攻击是一种试图诱骗人们泄露敏感信息的消息。现有的基于机器学习和黑名单的检测方法为了避免拦截合法消息,对新的攻击既脆弱又有些宽松;因此,广泛使用的电子邮件系统很容易受到精心设计的钓鱼电子邮件的攻击。为了解决这个问题,项目组将开发系统,自动阻止明显的诈骗,同时将较少的特定案例转发给接受过检测网络钓鱼邮件培训的群组工作人员。为了支持这些工人的决策,该团队将对系统的决策做出新的解释,突出信息及其算法的方面,这些方面引发了人类判断的需要。该系统还将聚合这些人群决策,以生成实时网络钓鱼警报,可以共享给个人用户和电子邮件系统。该项目将导致可解释机器学习方面的进展,鉴于人工智能和机器学习系统在社会中发挥的作用越来越大,这是一个重要的主题,还将提高我们表征钓鱼攻击的演变以及互联网平台和用户随着时间的推移对这些攻击的脆弱性的能力。该项目团队还将把这项工作作为可用安全的新课程的重要组成部分,并向高中教师和学生提供扩展计划,以教育他们并增加他们对网络安全研究的参与。工作围绕三个主要目标组织:对网络钓鱼风险进行经验表征,为网络钓鱼检测开发准确和可解释的机器学习模型,以及为网络钓鱼警报开发可靠的众包系统。该团队将通过开发电子邮件系统中反欺骗协议的有效采用和配置的分析工具来评估网络钓鱼风险,使用对抗性机器学习方法在现有网络钓鱼检测器上进行黑盒测试,并创建引诱和响应网络钓鱼攻击的反应性蜜罐,以便不仅收集初始网络钓鱼电子邮件的数据,还收集攻击者在整个网络钓鱼攻击过程中的行为数据。从钓鱼电子邮件中收集的数据将被用来开发机器学习模型,使用卷积神经网络和基于长短期记忆的深度学习技术来生成可疑特征和对个人决策的置信度估计。可疑特征将用于生成可解释的安全提示,例如文本注释或图标,方法是首先创建更简单、更可解释的机器学习模型,例如模仿特征空间中目标电子邮件附近的本地检测边界的决策树。决策树中的规则将被映射回界面元素和电子邮件内容以提供警告,并将这些警告与一系列用户研究中的通用电子邮件安全警告进行比较,这些研究还模拟了人们使用各种线索、功能和媒体检测网络钓鱼的能力。然后,这些单独的模型以及钓鱼检测模型的置信度估计将被用于驱动基于众包的系统,在该系统中,个人用户质量的模型将被聚合,以针对被评为太可疑而无法通过但又不够可疑而无法自动过滤的电子邮件做出可靠的判断。这一裁决反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project aims to address the growing threat of phishing attacks, messages that try to trick people into revealing sensitive information, by combining human and machine intelligence. Existing detection methods based on machine learning and blacklists are both brittle to new attacks and somewhat lenient, in order to avoid blocking legitimate messages; as a result, widely used email systems are vulnerable to carefully crafted phishing emails. To address this, the project team will develop systems that automatically block obvious scams while forwarding less certain cases to groups of crowd workers trained to detect phishing mails. To support these workers' decision-making, the team will develop novel explanations of the system's decision making that will highlight the aspects of both the message and its algorithm that triggered the need for human judgment. The system will also aggregate these crowd decisions to generate real-time phishing alerts that can be shared to both individual users and to email systems. The project will lead to advances in interpretable machine learning, an important topic given the increasing role that artificial intelligence and machine learning systems play in society, and also increase our ability to characterize the evolution of phishing attacks and the vulnerability of internet platforms and users to those attacks over time. The project team will also use the work as an important component of new courses on usable security and outreach programs to high school teachers and students to both educate them about and increase their participation in cybersecurity research.The work is organized around three main objectives: empirical characterization of phishing risks, developing accurate and interpretable machine learning models for phishing detection, and developing reliable crowdsourcing systems for phishing alerts. The team will assess phishing risks through developing analytics tools on the effective adoption and configuration of anti-spoofing protocols in email systems, using adversarial machine learning methods to conduct black box testing on existing phishing detectors, and creating reactive honeypots that entice and respond to phishing attacks in order to collect data on not just the initial phishing emails but on attackers' behaviors throughout the course of a successful phishing attack. The data collected on phishing emails will be used to develop the machine learning models, using Convolutional Neural Network and Long Short-Term Memory based deep learning techniques to generate both suspicious features and confidence estimates of individual decisions. The suspicious features will be used to generate interpretable security cues such as text annotations or icons by first creating simpler and more interpretable machine learning models such as decision trees that mimic the local detection boundary near the target emails in the feature space. Rules in the decision tree will be mapped back to interface elements and email content to provide the warnings, and these will be compared to generic email security warnings in a series of user studies that also model people's ability to detect phishing using a variety of cues, features, and media. Those individual models, along with the confidence estimates from the phishing detection model, will then be used to drive a crowdsourcing-based system where the models of individual users' quality will be aggregated to make reliable judgments around emails the models judge as too suspicious to pass but not suspicious enough to automatically filter.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2018
期刊:
影响因子:
--
作者:
[Dongliang Mu;A. Cuevas;Limin Yang;Hang Hu;Xinyu Xing;Bing Mao;G. Wang]
通讯作者:
Dongliang Mu;A. Cuevas;Limin Yang;Hang Hu;Xinyu Xing;Bing Mao;G. Wang
DOI:
--
发表时间:
2018
期刊:
影响因子:
--
作者:
[Hang Hu;G. Wang]
通讯作者:
Hang Hu;G. Wang
DOI:
10.1109/secdev.2018.00020
发表时间:
2018-09
期刊:
2018 IEEE Cybersecurity Development (SecDev)
影响因子:
--
作者:
[Hang Hu;Peng Peng-Peng;G. Wang]
通讯作者:
Hang Hu;Peng Peng-Peng;G. Wang
DOI:
10.1145/3355369.3355585
发表时间:
2019-10
期刊:
Proceedings of the Internet Measurement Conference
影响因子:
--
作者:
[Peng Peng-Peng;Limin Yang;Linhai Song;Gang Wang]
通讯作者:
Peng Peng-Peng;Limin Yang;Linhai Song;Gang Wang
DOI:
10.1145/3319535.3363195
发表时间:
2019-11
期刊:
Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
作者:
[Sazzadur Rahaman;Gang Wang;D. Yao]
通讯作者:
Sazzadur Rahaman;Gang Wang;D. Yao
共 12 条
Travel: NSF Student Travel Grant for the 2023 ACM International Conference on Mobile Systems, Applications, and Services (MobiSys)
-
批准号:2325485
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:2023
-
负责人:Gang Wang
-
依托单位:
Collaborative Research: SaTC: CORE: Small: Towards Label Enrichment and Refinement to Harden Learning-based Security Defenses
-
批准号:2055233
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2021
-
负责人:Gang Wang
-
依托单位:
SaTC: CORE: Small: Collaborative: Towards Facilitating Kernel Vulnerability Reproduction by Fusing Crowd and Machine Generated Data
-
批准号:1955719
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2020
-
负责人:Gang Wang
-
依托单位:
CAREER: Machine Learning Assisted Crowdsourcing for Phishing Defense
-
批准号:2030521
-
项目类别:Continuing Grant
-
资助金额:$42.04万
-
财政年份:2019
-
负责人:Gang Wang
-
依托单位:
Planning Grant: I/UCRC for Advanced Composites in Transportation Vehicles (ACTV)
-
批准号:1361904
-
项目类别:Standard Grant
-
资助金额:$1.15万
-
财政年份:2014
-
负责人:Gang Wang
-
依托单位:
国内基金
海外基金
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位: