Robust Estimation and Outlier Detection for Overdispersed Multinomial Models of Count Data

Robust Estimation and Outlier Detection for Overdispersed Multinomial Models of Count Data
复制标题

计数数据过分散多项式模型的鲁棒估计和异常值检测

DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
J. Sekhon
J. Sekhon
中科院分区:
--
文献类型:
--
作者:
W. Mebane;J. Sekhon

文献摘要

被引文献

相似文献

对于计数数据的过度离散多项回归模型,我们提出了一个稳健估计--双曲正切(TANH)估计。即使指定的模型不适用于多达一半的数据,TANH估计器也能提供准确的估计和可靠的推断。严重不符合的计数-离群值-被确定为估计的一部分。蒙特卡罗抽样实验表明,在实际样本量下,即使10%的数据是由显著不同的过程产生的,TANH估计器也能产生良好的结果。实验表明,在污染数据的情况下,使用其他四种估计量:非稳健极大似然估计量、加性Logistic模型和两个SUR模型,估计都是失败的。使用tanh估计器分析佛罗里达州2000年总统选举的数据,与其他四个估计器未能捕捉到的众所周知的选举特征相匹配。在对1993年波兰议会选举数据的分析中,tanh估计器给出的推论比之前提出的异方差sur模型更准确。
We develop a robust estimator—the hyperbolic tangent (tanh) estimator—for overdispersed multinomial regression models of count data. The tanh estimator provides accurate estimates and reliable inferences even when the specified model is not good for as much as half of the data. Seriously ill-fitted counts—outliers—are identified as part of the estimation. A Monte Carlo sampling experiment shows that the tanh estimator produces good results at practical sample sizes even when ten percent of the data are generated by a significantly different process. The experiment shows that, with contaminated data, estimation fails using four other estimators: the nonrobust maximum likelihood estimator, the additive logistic model and two SUR models. Using the tanh estimator to analyze data from Florida for the 2000 presidential election matches well-known features of the election that the other four estimators fail to capture. In an analysis of data from the 1993 Polish parliamentary election, the tanh estimator gives sharper inferences than does a previously proposed heteroskedastic SUR model.