Outlier Detection in High-Dimensional Big Data using Bio-Inspired Methods for Emerging Applications in Engineering, Healthcare, and Business
Outlier Detection in High-Dimensional Big Data using Bio-Inspired Methods for Emerging Applications in Engineering, Healthcare, and Business
批准号:
RGPIN-2017-04192
负责人:
Raahemi, Bijan
金额:
$1.75万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2020
资助国家:
加拿大
项目状态:
已结题
起止时间:
2020-01-01 至 2021-12-31
中文摘要
在这项研究计划中,我将探索,设计和分析创新算法,用于使用生物启发的方法在高维大数据中进行离群值检测,并将新方法应用于工程(即计算机网络中的入侵检测),商业(即公司财务报表中的欺诈检测)和医疗保健(即检测患者生命信号中的异常)中的新兴应用。
大数据的特征是四个(有时甚至更多)V,即数量、速度、多样性和准确性,它被定义为一组如此庞大、动态和复杂的数据集,以至于使用传统的数据分析技术很难处理。在高维空间中,点之间的距离变得相对均匀,数据点的最近邻居的概念变得毫无意义。 一个高维数据也有许多排列的子空间,这是实际上是不可行的,以检查所有。 处理这样的大规模多维数据在计算上是复杂且昂贵的。
检测大数据中的异常值(相对于大多数数据非常不相似和不一致的对象),特别是在高维数据和存在噪声的情况下,是一个重要的研究问题,由于其引入的科学挑战,以及它支持的广泛的现实世界应用,包括工程,医疗保健,商业,环境和公共安全,因此引起了研究界的广泛关注。
在这个研究项目中,我将探索在分布式平台上运行的降维,数据汇总和特征转换的新技术,结合模型集成来快速准确地检测离群值。我将探索生物启发的算法,以搜索一个大的空间的排列与健身功能,最大限度地减少稀疏的样本在选定的子空间。
大数据分析依赖于可扩展的分布式平台,如Hadoop(支持MapReduce结构,用于并行分析大数据)和Spark(用于大规模数据处理的快速内存引擎)。在我的知识发现和数据挖掘实验室中,我们已经尝试使用Hadoop和Spark并行处理任务。在这些经验的基础上,我们将在分布式平台上设计和实施我们的新解决方案。
在这个研究项目中发现的解决方案和算法将应用于3个工程领域的新兴应用(分析互联网流量产生的大量高维数据,以检测网络中的入侵),业务(在彭博社和CompuStat提供的4000多家公司的真实的数据集中检测金融欺诈活动),和医疗保健(分析从患者收集的生命信号,包括体温、心跳、血压和ECG信号,以检测异常)。
英文摘要
In this research program, I will explore, design, and analyze innovative algorithms for outlier detection in high-dimensional big data using bio-inspired approaches, and apply the new methods to emerging applications in engineering (namely, Intrusion detection in computer networks), business (namely, fraud detection in corporate financial statements), and healthcare (namely, detecting abnormalities in patient's vital signals).
Big data, characterized by four (and sometimes more) V's of Volume, Velocity, Variety, and Veracity, is defined as a collection of data sets so large, dynamic, and complex that it becomes difficult to process using traditional data analytics techniques. In high dimensional spaces, distances between points become relatively uniform, and the notion of the nearest neighbors of a data point becomes meaningless. A high dimensional data has also numerous permutations of sub-spaces which are practically infeasible to be examined all. Processing such large-scale multi-dimensional data is computationally complex and expensive.
Detecting outliers (objects considerably dissimilar and inconsistent with respect to the majority of data) in Big data, especially in high-dimensional data and in the presence of noise, is an important research problem which has drawn many attentions in research community due to scientific challenges it introduces, and a wide range of real-world applications it supports including in engineering, healthcare, business, environment, and public security.
In this research program, I will explore novel techniques for dimension reduction, data summarization, and feature transformation running on distributed platforms, combined with ensemble of models to make fast and accurate detection of outliers. I will explore bio-inspired algorithms to search a large space of permutations with fitness functions minimizing sparsity of the samples in selected sub-spaces.
Analysis of Big data relies on scalable distributed platforms such as Hadoop (which supports MapReduce structure for analysis of large data in parallel), and Spark (a fast in-memory engine for large scale data processing.). In my Knowledge Discovery and Data Mining Lab, we have experimented with processing tasks in parallel using Hadoop and Spark. Building on these experiences, we will design and implement our novel solutions on distributed platforms.
The solutions and algorithms discovered in this research program will be applied to emerging applications in 3 areas of engineering (analyzing large volume of high-dimensional data generated by Internet traffic to detect intrusion in the network), business (detecting financial fraudulent activities in a real dataset of more than 4000 firms provided by Bloomberg, and CompuStat), and healthcare (analyzing vital signals collected from patients including temperature, heartbeat, blood pressure, and ECG signals to detect anomalies).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Outlier Detection in High-Dimensional Big Data using Bio-Inspired Methods for Emerging Applications in Engineering, Healthcare, and Business
-
批准号:RGPIN-2017-04192
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$3.5万
-
财政年份:2021
-
负责人:Raahemi, Bijan
-
依托单位:
Outlier Detection in High-Dimensional Big Data using Bio-Inspired Methods for Emerging Applications in Engineering, Healthcare, and Business
-
批准号:RGPIN-2017-04192
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2019
-
负责人:Raahemi, Bijan
-
依托单位:
Outlier Detection in High-Dimensional Big Data using Bio-Inspired Methods for Emerging Applications in Engineering, Healthcare, and Business
-
批准号:RGPIN-2017-04192
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2018
-
负责人:Raahemi, Bijan
-
依托单位:
Outlier Detection in High-Dimensional Big Data using Bio-Inspired Methods for Emerging Applications in Engineering, Healthcare, and Business
-
批准号:RGPIN-2017-04192
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.75万
-
财政年份:2017
-
负责人:Raahemi, Bijan
-
依托单位:
Estimating Bus Passengers' Origin Destination Travel Route using Data Analytics on Wi-Fi and Bluetooth Signals
-
批准号:514854-2017
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2017
-
负责人:Raahemi, Bijan
-
依托单位:
Feature Engineering using Bio-Inspired Methods for the Internet Data Analytics
-
批准号:341811-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.53万
-
财政年份:2016
-
负责人:Raahemi, Bijan
-
依托单位:
Feature Engineering using Bio-Inspired Methods for the Internet Data Analytics
-
批准号:341811-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.53万
-
财政年份:2015
-
负责人:Raahemi, Bijan
-
依托单位:
Capturing and analyzing data from Giatec's testing devices using web applications
-
批准号:463717-2014
-
项目类别:Engage Grants Program
-
资助金额:$1.69万
-
财政年份:2014
-
负责人:Raahemi, Bijan
-
依托单位:
Capturing and analyzing data from SensorSuite's sevices using big data analytics techniques
-
批准号:477440-2014
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2014
-
负责人:Raahemi, Bijan
-
依托单位:
Feature Engineering using Bio-Inspired Methods for the Internet Data Analytics
-
批准号:341811-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.53万
-
财政年份:2014
-
负责人:Raahemi, Bijan
-
依托单位:
Feature Engineering using Bio-Inspired Methods for the Internet Data Analytics
-
批准号:341811-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.53万
-
财政年份:2013
-
负责人:Raahemi, Bijan
-
依托单位:
Feature Engineering using Bio-Inspired Methods for the Internet Data Analytics
-
批准号:341811-2012
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.53万
-
财政年份:2012
-
负责人:Raahemi, Bijan
-
依托单位:
Stream Data mining in telecommunication industry with privacy preserving consideration
-
批准号:341811-2007
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.09万
-
财政年份:2009
-
负责人:Raahemi, Bijan
-
依托单位:
Stream Data mining in telecommunication industry with privacy preserving consideration
-
批准号:341811-2007
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.09万
-
财政年份:2008
-
负责人:Raahemi, Bijan
-
依托单位:
Stream Data mining in telecommunication industry with privacy preserving consideration
-
批准号:341811-2007
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.09万
-
财政年份:2007
-
负责人:Raahemi, Bijan
-
依托单位:
国内基金
海外基金
Graphon mean field games with partial observation and application to failure detection in distributed systems
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2025
-
负责人:MATHIEULOUROCHLAURIERE
-
依托单位: