Distribution-Free Detection of Structured Anomalies: Permutation and Rank-Based Scans

Distribution-Free Detection of Structured Anomalies: Permutation and Rank-Based Scans
复制标题

DOI:
10.1080/01621459.2017.1286240
复制
发表时间:
2018-01-01
影响因子:
3.7
通讯作者:
Wang, Meng
Wang, Meng
中科院分区:
数学1区
文献类型:
--
作者:
Arias-Castro, Ery;Castro, Rui M.;Wang, Meng

文献摘要

被引文献

相似文献

扫描统计量是迄今为止最流行的异常检测方法,在综合征监视,信号和图像处理中流行,以及基于传感器网络的目标检测以及其他应用。在这种情况下,扫描统计数据的使用得出了一个假设测试程序,其中零假设对应于缺乏异常行为。如果已知无效分布,则基于扫描的测试的校准相对容易,因为它可以通过Monte Carlo Simulation完成。当零分布未知时,它的简单性就不那么直接了。我们研究了两个程序。第一个是按排列进行校准,另一个是基于等级的扫描测试,该测试无分布且对异常值较不敏感。此外,排名扫描测试仅需要一次性校准,对于给定的数据大小,使其在计算上更具吸引力。在这两种情况下,我们都量化了知道无效分布的Oracle扫描测试的性能损失。我们表明,使用这些校准程序之一,在自然指数家族的背景下,仅会导致非常小的功率损失。这包括在信号处理中流行的经典正常位置模型和在综合症监测中流行的Poisson模型。我们对模拟数据进行数值实验,进一步支持我们的理论以及基因组学的真实数据集。本文的补充材料可在线获得。
The scan statistic is by far the most popular method for anomaly detection, being popular in syndromic surveillance, signal and image processing, and target detection based on sensor networks, among other applications. The use of the scan statistics in such settings yields a hypothesis testing procedure, where the null hypothesis corresponds to the absence of anomalous behavior. If the null distribution is known, then calibration of a scan-based test is relatively easy, as it can be done by Monte Carlo simulation. When the null distribution is unknown, it is less straightforward. We investigate two procedures. The first one is a calibration by permutation and the other is a rank-based scan test, which is distribution-free and less sensitive to outliers. Furthermore, the rank scan test requires only a one-time calibration for a given data size making it computationally much more appealing. In both cases, we quantify the performance loss with respect to an oracle scan test that knows the null distribution. We show that using one of these calibration procedures results in only a very small loss of power in the context of a natural exponential family. This includes the classical normal location model, popular in signal processing, and the Poisson model, popular in syndromic surveillance. We perform numerical experiments on simulated data further supporting our theory and also on a real dataset from genomics. Supplementary materials for this article are available online.