Robust mislabel logistic regression without modeling mislabel probabilities

Robust mislabel logistic regression without modeling mislabel probabilities
复制标题

DOI:
10.1111/biom.12726
复制
发表时间:
2018-03-01
期刊:
影响因子:
1.9
通讯作者:
Huang, Su-Yun
Huang, Su-Yun
中科院分区:
数学3区
文献类型:
--
作者:
Hung, Hung;Jou, Zhi-Yu;Huang, Su-Yun

文献摘要

被引文献

相似文献

逻辑回归是线性判别分析中使用最广泛的统计方法之一。在许多应用中,我们只观察到可能贴错标签的响应。拟合传统的逻辑回归可能会导致估计有偏差。一种常见的解决方案是拟合错误标签的逻辑回归模型,该模型考虑了错误标签的响应。另一种常见方法是通过降低可疑实例的权重来采用稳健的 M 估计。在这项工作中,我们提出了一种基于散度的新的鲁棒错误标签逻辑回归。我们的建议具有两个有利特征:(1)它不需要对错误标签概率进行建模。 (2)最小散度估计得到的加权估计方程不需要包含任何偏差校正项,即自动进行偏差校正。这些特征使得所提出的逻辑回归在模型拟合中更加稳健,并且通过简单的加权方案使模型解释更加直观。我们的方法也很容易实现,并且包括两种类型的算法。提出了模拟研究和 Pima 数据应用来证明 -logistic 回归的性能。
Logistic regression is among the most widely used statistical methods for linear discriminant analysis. In many applications, we only observe possibly mislabeled responses. Fitting a conventional logistic regression can then lead to biased estimation. One common resolution is to fit a mislabel logistic regression model, which takes into consideration of mislabeled responses. Another common method is to adopt a robust M-estimation by down-weighting suspected instances. In this work, we propose a new robust mislabel logistic regression based on -divergence. Our proposal possesses two advantageous features: (1) It does not need to model the mislabel probabilities. (2) The minimum -divergence estimation leads to a weighted estimating equation without the need to include any bias correction term, that is, it is automatically bias-corrected. These features make the proposed -logistic regression more robust in model fitting and more intuitive for model interpretation through a simple weighting scheme. Our method is also easy to implement, and two types of algorithms are included. Simulation studies and the Pima data application are presented to demonstrate the performance of -logistic regression.