Outliers in Time Series

Outliers in Time Series
复制标题

DOI:
10.1111/j.2517-6161.1972.tb00912.x
复制
发表时间:
1972-07
期刊:
Journal of the royal statistical society series b-methodological
影响因子:
--
通讯作者:
J. Peter But-Man;M. Otto;Nash J. Monsour
J. Peter But-Man;M. Otto;Nash J. Monsour
中科院分区:
其他
文献类型:
--
作者:
J. Peter But-Man;M. Otto;Nash J. Monsour

文献摘要

被引文献

相似文献

尽管最近的一些工作也涉及标准线性模型,但异常值的检测主要考虑的是单个随机样本;例如,参见 Anscombe (1960) 和 Kruskal (1960)。本质上类似的问题出现在时间序列中(Burman,1965),但似乎没有发表的工作考虑到连续观察之间的相关性。过去,时间序列中异常值的搜索基于观测值独立且同正态分布的假设。这种假设导致的分析被称为随机抽样程序。本文考虑了时间序列中可能出现的两种异常值。 I 类异常值对应于观察的严重错误或记录错误影响单个观察的情况。第二类异常值对应于单个“创新”极端的情况。这不仅会影响特定的观察,还会影响后续的观察。为了开发测试和解释异常值,有必要区分过程中可能包含的异常值类型。目前的方法基于该问题的四种可能的表述:异常值都是 I 类;异常值均为II类;异常值都属于同一类型,但不知道它们是 I 型还是 II 型;异常值是两种类型的混合。由于比似然比方法给出的解决方案更实用的解决方案通常是从似然比准则的简化中获得的,因此导出了一些更简单的准则。这些标准的形式为 /&2a,其中 A 是测试观测值的估计误差,^ 是 A 的估计标准误差。在本文中,趋势和季节性分量被假定为可忽略不计或已被消除。去除这些成分所采用的方法可能会以某种方式影响结果。
THE detection of outliers has mainly been considered for single random samples, although some recent work deals also with standard linear models; see, for example, Anscombe (1960) and Kruskal (1960). Essentially similar problems arise in time series (Burman, 1965) but there seems no published work taking into account correlations between successive observations. In the past, the search for outliers in time series has been based on the assumption that the observations are independently and identically normally distributed. This assumption leads to analyses which will be called random sample procedures. Two types of outlier that may occur in a time series are considered in this paper. A Type I outlier corresponds to the situation in which a gross error of observation or recording error affects a single observation. A Type II outlier corresponds to the situation in which a single "innovation" is extreme. This will affect not only the particular observation but also subsequent observations. For the development of tests and the interpretation of outliers, it is necessary to distinguish among the types of outlier likely to be contained in the process. The present approach is based on four possible formulations of the problem: the outliers are all of Type I; the outliers are all of Type II; the outliers are all of the same type but whether they are of Type I or of Type II is not known; and the outliers are a mixture of the two types. Since more practical solutions than those given by likelihood ratio methods are often obtained from simplifications of likelihood ratio criteria, some simpler criteria are derived. These criteria are of the form /&2a, where A is the estimated error in the observation tested and ^ is the estimated standard error of A. Throughout this paper, trend and seasonal components are assumed either negligible or to have been eliminated. The method adopted to remove these components might affect the results in some way.