Expanding and Assembling Approaches to Improve Decisions on Identification and Classification of Online Terrorist Content
Expanding and Assembling Approaches to Improve Decisions on Identification and Classification of Online Terrorist Content
批准号:
2440634
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
迫切需要加强对网上恐怖主义生态系统的了解,以便有效和负责任地制定反措施。该项目的目的是开发新的方法来预测互联网上的恐怖主义和极端主义行为。行为模式的准确模型将使人们能够预测通信渠道,并通过扩大已部署的技术方法(包括管理发现/开源情报和通过机器学习与平台互动以改进决策)和确定将各种方法组合成整体模型的机会,使内容发现(和删除)战略更加有效。这个博士项目旨在应用最先进的技术和新的算法技术来分类与极端主义相关的数据和行为,方法是开发以人为中心的流程,从将从一系列社交媒体平台收集的数据中消除偏见和噪音。为了实现我们的目标,首先,我们将使用TCAP平台,该平台将为我们提供访问以不同媒体形式(PDF、URL、HTML、音频和视频)从一系列社交媒体平台收集数据的支持。原始数据将被转换为可使用的数据集,如SQLite或CSV文件格式。其次,为了消除数据中的噪音和偏差,我们将促进以用户为中心的设计过程,在该过程中,我们将开发一个交互过程,使极端主义领域的专家能够大规模执行复杂的文本提取任务,如[5,7]所述。该工具将使用户能够以可量化的方式消除噪音,从而使我们能够挤压清洁过程和用户之间的反馈回路。第三,我们将应用新的算法技术对与极端主义有关的数据和行为进行分类。这将包括集合方法和/或深度学习技术。一种方法可以基于基于情感的深度学习模型(LSTM CNN)来对极端主义和非极端主义内容进行分类[8]。我们将通过以下方式将该过程应用于一组封闭的数据。我们将重点分析过去恐怖组织利用不同平台传播宣传的事件。该项目将选择最近的一次历史性恐怖袭击,并将分析恐怖分子在这四个袭击阶段使用社交媒体的情况。在这里,我们可以使用攻击的四个不同时间段内的现有数据(来自TCAP)。在第一步中,学生将围绕四个阶段的用户交互生成一个数据集。在第二阶段,将遵循以用户为中心的交互式流程,通过让领域专家参与来帮助清理数据。请注意,在此过程中,我们将考虑避免过度贴合的因素。在第三步,也是最后一步,我们将应用最先进的机器学习模型,如基于情感的深度学习,和/或集成模型来开发用于识别极端主义内容类型的分类器。将在主要里程碑进行用户研究,以对机器学习模型提供的决策支持进行评级,并确保决策是合理的,不违反言论自由原则。一种交互式的、以用户为中心的工具,使极端主义领域的专家能够协助进行数据清理和消除偏见。一系列分类器,用于预测数据集上不同类型的极端主义内容的类别,这些内容将公开提供给未来的研究。该项目将调查分类的有效性,以限制侧重于恐怖主义活动和宣传的内容。开发的方法和技术将受到仔细审查,以确定其可能的意外用途,如对民主的攻击,对政治决策的不当影响,以及其他欺诈性行为,或故意将偏见引入社交媒体版图。
英文摘要
Improving knowledge of online terrorist ecosystems is urgently needed to develop counter measures effectively and responsibly. The aim of this project is to develop new methods for predicting terrorist and extremist behaviour on the Internet. Accurate models of behavioural patterns will allow the predictionof communication channels and make content discovery (and removal) strategies more effective by expanding technological approaches deployed (including managing discovery- /open-source intelligence and interacting with platforms to improve decisions via machine learning) and scoping opportunities to combine approaches into ensemble models, such as combining predictive algorithms. This PhD project aims to apply state of the art and novel algorithmic techniques to classify data, and behaviours related to extremism by developing human-centred processes to clean bias and noise from the data that will be collected from a range of social media platforms. To achieve our aims, firstly, we will use the TCAP platform that will provide us support on getting access to collecting data from a range of social media platform in different media forms (PDF, URL,HTML, audio and videos). The raw data will be converted into a workable dataset such as SQLite, or csv file format. Secondly, to remove noise and bias from the data, we will facilitate a user-centred design process in which we will develop an interactive process that will enable extremist domain experts to perform complex text extraction tasks at scale, as described in [5,7]. The tool will enable users to remove noise in quantifiable ways which will consequently allow us to squeeze the feedback loop between the cleaning process and the user. Thirdly, we will apply novel algorithmic techniques to classify data, and behaviours related to extremism. This will include ensemble methods and/or deep learning techniques. An approach can be based on sentiment-based deep learning models (LSTM+CNN) to classify extremist and non-extremist content [8]. We will apply the process on a closed set of data in the following way. We will focus on analyzing the past occurrences where terrorist organisations have used different platforms to spread propaganda. This project will select one of the recent historical terrorist attacks and will analyze the use of social media by terrorist across these four stages of the attacks. Here, we can use the existing data (from TCAP) during the four different time frames of an attack. In the first step, the student will generate a dataset around user interactions during the four stages. In the second stage, an interactive user centered process will be followed that will help clean the data by involving domain experts. Note, we will consider factors to avoid overfitting during this process. In the third and final step, we will apply state-of-the-art machine learning models such as sentiment-based deep learning, and/or ensemble models to develop classifiers for identifying the type of extremist contents. User-studies will be conducted at major milestones to rate the decisionsupport provided by the machine learning models and to ensure the decisions are justifiable and do not violate the principle of freedom of speech.In summary, the following contributions are anticipated. An interactive, user-centered tool that enables extremism domain experts to assist with data cleaning and removing bias. A range of classifiers to predict the class of different type of extremist contents on a dataset which will be made available publicly to inform future research. The project will investigate the effectiveness of classification for the purpose of content restriction focused on terrorist activity and propaganda. The methodology and technology developed will be scrutinised for its possible unintended use, such attacks on democracy, undue influencing of political decisions, and other fraudulent behaviour, or deliberate introduction of biases into the social media landscape.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金