II-NEW: Collaborative Research: Spam Processing, Archiving, and Monitoring Community Facility (SPAM Commons)
II-NEW: Collaborative Research: Spam Processing, Archiving, and Monitoring Community Facility (SPAM Commons)
批准号:
1118355
负责人:
Brent Kang
金额:
$5.04万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-12-01 至 2012-08-31
中文摘要
在这个项目中,pi建议构建和开发一个共享的基础设施,以支持实际的、大规模的垃圾邮件数据集的收集和维护,称为spam Commons。垃圾邮件是许多重要的通信媒体,如电子邮件和网络的一个问题。垃圾邮件的一个子问题,网络钓鱼(一种在线借口),在2007年造成了大约32亿美元的损失。有效的垃圾邮件过滤方法的广泛影响可以估计在几个通信媒体,如电子邮件和网络数十亿美元。垃圾邮件也侵入了其他媒体,具体的攻击例子包括社交网络、博客圈、互联网电话(VoIP)、即时消息和点击欺诈。不幸的是,由于对隐私和公司知识产权的担忧,缺乏公开的真实世界数据集,阻碍了垃圾邮件的研究。这个项目团队开发了一个共享的基础设施来支持实际的、大规模的垃圾邮件数据集的收集和维护,称为垃圾邮件处理、归档和监控社区设施(spam Commons)。垃圾邮件共享的主要目标是:(1)促进补救研究,以阻止垃圾邮件造成的浪费和损失,(2)使旨在完全阻止某些类型的垃圾邮件攻击的革命性研究成为可能。SPAM Commons分为公共分区和保护分区。公共分区是语音和图像识别研究的标准语料库的直接模拟,由各种通信媒体中垃圾邮件和合法数据的系统和定期收集组成,从电子邮件和网络垃圾邮件开始,随着垃圾邮件在每个领域成为严重威胁并且数据可用,扩展到其他通信媒体。受保护分区由数据和处理设施组合而成,该设施使私有数据或接近实时的垃圾邮件数据可用于在受保护的测试平台中对垃圾邮件防御机制进行实验性评估。访问这些受保护的数据将使新的垃圾邮件研究能够实时发展的垃圾邮件和现实世界的数据集,这在今天是不可行的。SPAM Commons项目的智力挑战超出了对上述各种垃圾邮件领域的新研究,这些领域是由数据集的可用性支持的。SPAM Commons的这两个分区的构建本身就面临着重大的智力挑战。首先,保护分区的隔离部分地解决了隐私问题,这仍然是一个普遍的研究问题。其次,有用的垃圾邮件和合法数据集需要自动区分垃圾邮件和合法文档,这在电子邮件、web和其他媒体中仍然是一个开放的研究问题。第三,垃圾邮件制造者和捍卫者的对抗和相互演变需要不断收集新数据以供进一步研究。最后,近乎实时的垃圾邮件数据的收集和流代表了垃圾邮件研究人员目前无法获得的研究资源。这些领域的进步将刺激垃圾邮件共享的发展和演变,这将使对不断发展和增长的垃圾邮件问题的新研究成为可能。SPAM Commons数据集对实验性垃圾邮件研究的影响可能类似于语音/图像识别和自然语言处理等学科中大型语料库的影响,在使用此类语料库成为标准要求后,这些学科实现了一定程度的科学结果可重复性和可比性。拟议的数据存储库将由9所大学合作伙伴(克莱顿州立大学、埃默里大学、佐治亚理工学院、北卡罗来纳农工大学、西北大学、德克萨斯农工大学、加州大学戴维斯分校、乔治亚大学、北卡罗来纳大学夏洛特分校)和几个行业合作伙伴(IBM、PureWire、Secure Computing)支持和使用。
英文摘要
In this project, the PIs propose to construct and develop a shared infrastructure to support the collection and maintenance of realistic, large scale spam data sets, referred as SPAM Commons.Spam is a problem in many important communications media such as email and web. A sub-problem of spam, phishing (a form of online pretexting), caused an estimated $3.2B in damages in 2007. The broad impact of effective spam filtering methods can be estimated in billions of dollars in several communications media such as email and web.Spam has also invaded other media, with concrete attack examples in social networks, blogosphere, Internet telephony (VoIP), instant messaging, and click fraud. Unfortunately, spam research has been hampered by the lack of published real world data sets due to concerns with privacy and company intellectual property. This project team develops a shared infrastructure to support the collection and maintenance of realistic, large scale spam data sets, called Spam Processing, Archiving, and Monitoring Community Facility (SPAM Commons). The main goals of SPAM Commons are: (1) to facilitate remedial research that will stem the wastes and losses caused by spam, and (2) enable revolutionary research that aim for stopping certain kinds of spam attacks altogether. SPAM Commons is divided into a Public Partition and a Protected Partition.The Public Partition is a direct analog of standard corpora for speech and image recognition research, consisting of a systematic and regular collection of both spam and legitimate data in the various communications media, starting from email and web spam, and expanding into other communications media as spam becomes a serious threat in each area and data become available. The Protected Partition consists of a combined data and processing facility that makes private data or near real-time spam data available for experimental evaluation of spam defense mechanisms in a protected testbed. Access to such protected data will enable new spam research on real-time evolving spam and real world data sets that is infeasible today. The intellectual challenges of the SPAM Commons project extend beyond the new research on various abovementioned spam areas enabled by the availability of data sets. The construction of both partitions of SPAM Commons includes significant intellectual challenges of their own. First, the isolation of Protected Partition addresses partially the concerns of privacy, which remains a general research problem. Second, useful spam and legitimate data sets require automated distinction of spam from legitimate documents with certainty, which remains an open research question in email, web, and other media. Third, the adversarial and mutual evolution of spam producers and defenders require continuous collection of fresh data for further study. Finally, the collection and streaming of near-real-time spam data represent research resources currently unavailable to spam researchers. Advances in these areas will spur the growth and evolution of SPAM Commons that will enable new research on the evolving and growing spam problem.The impact of SPAM Commons data sets on experimental spam research may be similar to the impact of large corpora in disciplines such as speech/image recognition and natural language processing, which achieved a level of scientific result reproducibility and comparativeness after the use of such corpora became standard requirements. The proposed data repository will be supported and used by 9 university partners (Clayton State, Emory, Georgia Tech, NC A&T, Northwestern, Texas A&M, UC Davis, U. Georgia, UNC Charlotte), and several industry partners (IBM, PureWire, Secure Computing).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Hands-on exercises on DETER testbed for security education
-
批准号:1118424
-
项目类别:Standard Grant
-
资助金额:$4.93万
-
财政年份:2010
-
负责人:Brent Kang
-
依托单位:
Collaborative Research: Hands-on exercises on DETER testbed for security education
-
批准号:0920179
-
项目类别:Standard Grant
-
资助金额:$9.86万
-
财政年份:2009
-
负责人:Brent Kang
-
依托单位:
II-NEW: Collaborative Research: Spam Processing, Archiving, and Monitoring Community Facility (SPAM Commons)
-
批准号:0855067
-
项目类别:Standard Grant
-
资助金额:$12.0万
-
财政年份:2009
-
负责人:Brent Kang
-
依托单位:
Collaborative Project: Focused Faculty Development Workshop on Cyber Games and Interactive Simulations
-
批准号:0723808
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Brent Kang
-
依托单位:
海外基金