Anomaly Detection and Modeling of Trajectories

Anomaly Detection and Modeling of Trajectories
复制标题

DOI:
--
复制
发表时间:
2012-08
影响因子:
2.4
通讯作者:
Junier B. Oliva
Junier B. Oliva
中科院分区:
工程技术3区
文献类型:
--
作者:
Junier B. Oliva

文献摘要

被引文献

相似文献

翻译后摘要:最近的繁荣,地理定位技术的可用性和使用创造了一个很大的需要了解轨迹数据集。然而,轨迹有几个内在的属性,使他们难以分析。首先,它们的时间序列性质使得应用传统技术具有挑战性。其次,大多数数据集包含许多点的轨迹,这导致了高维建模问题。第三,有几个相互竞争的概念的相似性/差异的轨迹。为了应对这些挑战,本文提出了几种使用统计和机器学习的方法,这些方法提供了对轨迹数据集的深入理解。特别是,本文提出了方法来执行异常检测和密度估计,并创建空间图形模型。提出了一种使用支持向量机(SVMs)和轨迹的各种空间表示以无监督方式检测数据集中的异常轨迹的技术。本文还重点介绍了密度估计技术,即为数据集中的每个轨迹提供一个可能性。为了有效地执行密度估计的轨迹,一个马尔可夫假设的独立性的下一个位置的轨迹给定其先前的位置和核密度估计(KDE)的组合进行了探索。最后,论文探讨了空间图形模型。无向图模型详细描述了一组随机变量的条件独立结构。给定稀疏性假设,这个概念用于为具有与之相关联的空间位置的指标变量构建图形模型,指示代理是否接近相应的位置。实验使用两个真实世界的数据集进行:自动识别系统跟踪的英吉利海峡船只和1949年至2011年大西洋热带风暴和飓风路径。
Abstract : The recent boom in the availability and use of geolocation technologies has created a great need to understand datasets of trajectories. However, trajectories have several intrinsic attributes that make them difficult to analyze. First, their time-series nature makes applying traditional techniques challenging. Second, most datasets contain trajectories of many points, making for a high-dimensional modeling problem. Third, there are several competing notions of similarity/difference in trajectories. To deal with these challenges, this thesis proposes several methods using statistics and machine learning that provide a deep understanding of trajectory datasets. In particular, the thesis brings forth methods to perform anomaly detection and density estimation and to create spatial graphical models. A technique is presented for detecting anomalous trajectories in a dataset in an unsupervised fashion using support vector machines (SVMs) and various spatial representations of trajectories. The thesis also focuses on techniques for density estimation, that is providing a likelihood for each trajectory in a dataset. To effectively perform density estimation on trajectories, a combination of a Markovian assumption on the independence of the next position of a trajectory given its previous positions and kernel density estimation (KDE) is explored. Lastly, the thesis explores spatial graphical models. Undirected graphical models detail the conditional independence structure of a set of random variables. Given sparsity assumptions, this concept is used to build graphical models for indicator variables that have spatial locations associated with them, indicating if an agent has come near the corresponding location. Experiments were run using two real-world datasets: Automatic Identification System-tracked shipping vessels in the English Channel and every Atlantic Ocean tropical storm and hurricane track from 1949 to 2011.