Outliers in multilevel data

Outliers in multilevel data
复制标题

DOI:
10.1111/1467-985x.00094
复制
发表时间:
1998-01-01
影响因子:
2
通讯作者:
Lewis, T
Lewis, T
中科院分区:
数学4区
文献类型:
--
作者:
Langford, IH;Lewis, T

文献摘要

被引文献

相似文献

本文为数据分析人员提供了一系列处理多水平数据中异常值的实用程序。它首先开发了几种技术的数据探索离群值和离群值分析,然后将这些应用到两个大规模的多层次数据集的教育背景下的离群值的详细分析。这些技术包括使用偏差减少,基于残差的措施,杠杆值,层次聚类分析和称为DFITS的措施。离群值分析在多水平数据集中比在单变量样本或一组回归数据中更复杂,其中离群值的概念很简单。在多层次的情况下,人们必须考虑,例如,在什么水平o(-)水平一个特定的反应是ou!此外,在一个层次上对某一特定答复的处理可能会影响该答复的状况或模型中其他层次上其他单位的状况。
This paper offers the data analyst a range of practical procedures for dealing with outliers in multilevel data. It first develops several techniques for data exploration for outliers and outlier analysis and then applies these to the detailed analysis of outliers in two large scale multilevel data sets from educational contexts. The techniques include the use of deviance reduction, measures based on residuals, leverage values, hierarchical cluster analysis and a measure called DFITS. Outlier analysis is more complex in a multilevel data set than in, say, a univariate sample or a set of regression data, where the concept of an outlying value is straightforward. In the multilevel situation one has to consider, for example, at what level o(-) levels a particular response is ou!lying, and in respect of which explanatory variables; furthermore, the treatment of a particular response at one level may affect its status or the status of other units at other levels in the model.