Outside the box: an alternative data analytics framework

Outside the box: an alternative data analytics framework
复制标题

DOI:
10.14313/jamris_2-2014/16
复制
发表时间:
2014-04
期刊:
J. Autom. Mob. Robotics Intell. Syst.
影响因子:
--
通讯作者:
P. Angelov
P. Angelov
中科院分区:
其他
文献类型:
--
作者:
P. Angelov

文献摘要

被引文献

相似文献

在本文中,提出了一种替代的数据分析框架,它是基于空间感知的概念,偏心率和典型性,代表密度和接近的数据空间。这种方法是统计的,但不同于传统的概率论,这是频率论的性质。它也不同于基于信念和可能性的方法以及确定性的第一原理方法,尽管它可以被视为确定性的,因为它为相同的数据提供了完全相同的结果。它也不同于主观的专家为基础的方法,如模糊集。它可以用来检测异常,故障,形成集群,类,预测模型,控制器。引入新的基于典型性和偏心性的数据分析(TEDA)的主要动机是这样一个事实,即数据分析感兴趣的真实的过程,例如气候、经济和金融、机电、生物、社会和心理等,通常是复杂的、不确定的和鲜为人知的,但不是纯粹随机的。与纯粹的随机过程不同,如掷骰子、掷硬币、从碗中选择彩球和其他游戏,真实的生活过程确实违反了传统概率论所要求的主要假设。同时,它们很少是确定性的(更确切地说,总是具有不确定性/噪声分量,这是不确定的),创建基于专家和信念的可能性模型是繁琐和主观的。尽管如此,不同的研究人员和从业者群体喜欢并使用上述方法之一,概率论(也许)是最广泛使用的方法之一。建议的新框架泰达是一个系统的方法,不需要事先假设,并可用于发展的异常和故障检测,图像处理,聚类,分类,预测,控制,过滤,回归等方法的范围在本文中,由于空间的限制,只有一些说明性的例子提供了旨在证明的概念。
In this paper, an alternative framework for data analytics is proposed which is based on the spatially-aware concepts of eccentricity and typicality which represent the density and proximity in the data space. This approach is statistical, but differs from the traditional probability theory which is frequentist in nature. It also differs from the belief and possibility-based approaches as well as from the deterministic first principles approaches, although it can be seen as deterministic in the sense that it provides exactly the same result for the same data. It also differs from the subjective expert-based approaches such as fuzzy sets. It can be used to detect anomalies, faults, form clusters, classes, predictive models, controllers. The main motivation for introducing the new typicality- and eccentricity-based data analytics (TEDA) is the fact that real processes which are of interest for data analytics, such as climate, economic and financial, electro-mechanical, biological, social and psychological etc., are often complex, uncertain and poorly known, but not purely random. Unlike, purely random processes, such as throwing dices, tossing coins, choosing coloured balls from bowls and other games, real life processes of interest do violate the main assumptions which the traditional probability theory requires. At the same time they are seldom deterministic (more precisely, have always uncertainty/noise component which is nondeterministic), creating expert and belief-based possibilistic models is cumbersome and subjective. Despite this, different groups of researchers and practitioners favour and do use one of the above approaches with probability theory being (perhaps) the most widely used one. The proposed new framework TEDA is a systematic methodology which does not require prior assumptions and can be used for development of a range of methods for anomalies and fault detection, image processing, clustering, classification, prediction, control, filtering, regression, etc. In this paper due to the space limitations, only few illustrative examples are provided aiming proof of concept.