课题基金 / 基金详情

Dealing with Extreme Class Imbalance Learning in Defense and Security Applications

Dealing with Extreme Class Imbalance Learning in Defense and Security Applications
处理国防和安全应用中的极端类别不平衡学习
批准号:
RGPIN-2014-04889
负责人:
Japkowicz, Nathalie
金额:
$1.89万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2014
资助国家:
加拿大
项目状态:
已结题
起止时间:
2014-01-01 至 2015-12-31

项目摘要

项目成果

Japkowicz, Nathalie的其他基金

相似基金

相关文献

中文摘要
翻译
国防和安全应用程序,例如危险大气排放、水下水雷或计算机网络攻击的检测等威胁监控,都必须处理相同的基础问题:描述需要检测的事件的数据极度稀缺和不一致。这给机器学习应用带来了巨大的挑战。抽样方法的目的是增加事件数据量,但在只有少数事件数据可用的极端情况下,目前这类方法尚未成功。国防和安全社区的方法通常是基于领域知识生成模拟数据,以‘补充’缺失的事件数据。然而,考虑到模拟模型通常的简单性,这种类型的响应通常不能令人满意。另一种机器学习的反应是使用单类学习(离群值检测)方法来对通常非常丰富的背景数据进行建模,并在遇到“可疑”离群值时发送信号。然而,目前这类方法通常远不如它们的二进制类方法强大。我在这份Discovery Grant申请中提出的研究计划将有效地解决极端的班级失衡问题。特别是,我提出了一种新的方法,称为消极学习。负面学习包括认识到,尽管我们没有足够的异常/威胁数据实例,但我们有许多正常/背景数据的实例。基于这一事实,我们认为异常类是从正常类中缺失的任何和所有实例,并且我们建议从这个“正常类的否定”中适当地抽样。这项技术将基于由四个步骤组成的以下生成性方法。在第一步中,将从背景数据推导出概率密度函数。在第二步中,将在该模型的低概率区域中人工生成新的事件数据,因为这些区域中的数据可以被认为是异常类的边界实例。这意味着,与合成少数过采样技术(SMOTE)等简单方法不同,合成少数过采样技术只在可用数据描述的凸壳内生成新样本,而我建议生成的数据远远超出可用数据的凸壳。第三步将包括利用领域知识和用户指南精炼第一步和第二步产生的人工数据集。特别是,我将使用主动学习及其导数以及领域知识集成来用前两个步骤没有生成的数据来增加异常类,并通过消除冗余或不可信的数据来裁剪它。这一过程的最后一步将是对新生成的数据集应用二进制分类器,并与现有的其他技术(基于二进制和一类)比较,对整个系统进行评价。这项研究将使机器学习技术应用于该领域遇到的问题变得更加现实,因此应该对国防和安全产生重大影响。为此目的开发的技术还将在其他领域中找到应用,例如在医学领域和文本挖掘中。这项研究将允许两名博士生和三名硕士生在我的指导下学习,并自始至终进行他们的学习。
英文摘要
Defense and Security applications, such as threat monitoring, e.g., the detection of hazardous atmospheric emissions, underwater mines, or computer network attacks, must all deal with the same underpinning problem: the extreme scarcity and disparity of data describing an event in need of detection. This creates a significant challenge for machine learning applications. Sampling methods have aimed to increase the amount of event data, but in extreme cases in which only a handful of event data are available, current methods of this kind have not yet been successful. The defense and security community's approach has typically been to generate simulated data, based on domain knowledge, to 'complement' the missing event data. However, given the usual simplicity of the simulation models, this type of response is generally unsatisfactory. Another machine learning response has been to use one-class learning (outlier detection) approaches to model background data, which is typically plentiful, and to send a signal when a 'suspected' outlier is encountered. However, current approaches of this kind are typically much less powerful than their binary-class counterparts. The research program that I propose in this Discovery Grant application will effectively address the extreme class imbalance problem. In particular, I propose a new approach called Negative Learning. Negative Learning consists of recognizing that although we do not have enough instances of abnormal/threat data, we have many instances of normal/background ones. Based on this fact, we consider the abnormal class to be any and all instances missing from the normal class, and we propose to appropriately sample from this “negative of the normal class”. The technique will be based on the following generative approach composed of four steps. In the first step, a probability density function will be derived from the background data. In the second step, new event data will be artificially generated in the low probability regions of that model, because data in those regions can be thought of as borderline instances of the abnormal class. This means that unlike in simple methods like the Synthetic Minority Over-sampling Technique (SMOTE) approach, which only generates new samples within the convex hull described by the available data, I propose to generate data far beyond the convex hull of the available ones. The third step will consist of refining the artificial data set produced by the first and second step using domain knowledge and user guidance. In particular, I will use active learning and its derivatives as well as domain knowledge integration to augment the abnormal class with data not generated by the first two steps, and to trim it by eliminating redundant or implausible data. The last step of the process will be the application of binary classifiers to the newly generated data set and the evaluation of the overall system as compared to other techniques currently available (binary- and one-class based). This research should have a significant impact on Defense and Security because it will make the application of Machine Learning techniques to the problems encountered in the field much more realistic. The techniques developed for this purpose will also find applications in other fields such as in the medical domain and in text mining. This research will allow two Ph.D. and three Master’s student to study under my supervision and carry out their studies from beginning to end.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Dealing with Extreme Class Imbalance Learning in Defense and Security Applications
  • 批准号:
    RGPIN-2014-04889
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $0.49万
  • 财政年份:
    2016
  • 负责人:
    Japkowicz, Nathalie
  • 依托单位:
Dealing with Extreme Class Imbalance Learning in Defense and Security Applications
  • 批准号:
    RGPIN-2014-04889
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2015
  • 负责人:
    Japkowicz, Nathalie
  • 依托单位:
Predicting traffic safety based on weather events
  • 批准号:
    484326-2015
  • 项目类别:
    Engage Grants Program
  • 资助金额:
    $1.82万
  • 财政年份:
    2015
  • 负责人:
    Japkowicz, Nathalie
  • 依托单位:
Predicting network failures using anomaly detection methods
  • 批准号:
    485098-2015
  • 项目类别:
    Engage Grants Program
  • 资助金额:
    $1.82万
  • 财政年份:
    2015
  • 负责人:
    Japkowicz, Nathalie
  • 依托单位:
海外基金