Nonparametric Detection of Anomalous Data Streams

Nonparametric Detection of Anomalous Data Streams
复制标题

DOI:
10.1109/tsp.2017.2733472
复制
发表时间:
2014-04
影响因子:
5.4
通讯作者:
Shaofeng Zou;Yingbin Liang;H. V. Poor;Xinghua Shi
Shaofeng Zou;Yingbin Liang;H. V. Poor;Xinghua Shi
中科院分区:
工程技术1区
文献类型:
--
作者:
Shaofeng Zou;Yingbin Liang;H. V. Poor;Xinghua Shi

文献摘要

被引文献

相似文献

研究了一个非参数异常假设检验问题,其中共有$n$个观测序列,需要从其中检测出$s$个异常序列.每个典型序列由$m$独立同分布(i.i.d.)样本从分布$p$,而每个异常序列由$m$ i.i.d.从不同于$p$的分布$q$中提取的样本。假设分布$p$和$q$事先是未知的。分布自由的测试是通过使用最大的平均差异作为度量,这是基于平均嵌入到再生核希尔伯特空间的分布。错误的概率是有界的样本大小$m$,异常序列的数量$s$,和序列的数量$n$的函数。它表明,与$s$已知,构造的测试是指数一致的,如果$m$是大于一个常数因子$\log n$,任何$p$和$q$,而与$s$未知,$m$应该有一个顺序严格大于$\log n$。此外,它表明,没有测试可以是一致的任意$p$和$q$,如果$m$是小于一个常数因子$\log n$。因此,建立了所提出的测试的顺序级最优性。数值结果表明,所提出的测试优于(或执行以及)测试的基础上其他竞争的方法在各种情况下。
A nonparametric anomalous hypothesis testing problem is investigated, in which there are totally $n$ observed sequences out of which $s$ anomalous sequences are to be detected. Each typical sequence consists of $m$ independent and identically distributed (i.i.d.) samples drawn from a distribution $p$ , whereas each anomalous sequence consists of $m$ i.i.d. samples drawn from a distribution $q$ that is distinct from $p$ . The distributions $p$ and $q$ are assumed to be unknown in advance. Distribution-free tests are constructed by using the maximum mean discrepancy as the metric, which is based on mean embeddings of distributions into a reproducing kernel Hilbert space. The probability of error is bounded as a function of the sample size $m$, the number $s$ of anomalous sequences, and the number $n$ of sequences. It is shown that with $s$ known, the constructed test is exponentially consistent if $m$ is greater than a constant factor of $\log n$, for any $p$ and $q$ , whereas with $s$ unknown, $m$ should have an order strictly greater than $\log n$. Furthermore, it is shown that no test can be consistent for arbitrary $p$ and $q$ if $m$ is less than a constant factor of $\log n$. Thus, the order-level optimality of the proposed test is established. Numerical results are provided to demonstrate that the proposed tests outperform (or perform as well as) tests based on other competitive approaches under various cases.