Model selection via testing: an alternative to (penalized) maximum likelihood estimators

Model selection via testing: an alternative to (penalized) maximum likelihood estimators
复制标题

DOI:
10.1016/j.anihpb.2005.04.004
复制
发表时间:
2006-05
影响因子:
1.5
通讯作者:
L. Birgé
L. Birgé
中科院分区:
数学2区
文献类型:
--
作者:
L. Birgé

文献摘要

被引文献

相似文献

本文定义并研究了一类面向模型选择的估计量,我们称之为T-估计量(T表示检验)。他们的建设是基于以前的想法推导估计从一些家庭的测试,由于乐凸轮[LM乐凸轮,收敛的估计下维数限制,安。1(1973)38-53和LM Le Cam,关于实验渐近正态性理论中的局部和全局性质,在:M. Puri(Ed.),《随机过程与相关主题》,第1卷,学术出版社,纽约,1975年,第100页。13-54]和Birgé [L. Birgé,Approximation dans les espaces metriques et théorie de l'estimation,Z. Wahrscheinlichkeitstheorie Verw. Gebiete 65(1983)181-237,L. Birgé,Sur un théorème de minimax et son application aux tests,Probab.数学统计学家3(1984)259-282和L. Birgé,Stabilité et instabilité du risque minimax pour des variables indépendantes équidistribuées,Ann. Inst. H. Poincaré Sect. B 20(1984)201-223]以及关于来自巴伦和Cover的基于复杂度的模型选择[AR巴伦,TM Cover,最小复杂度密度估计,IEEE Trans. Inform.理论37(1991)1034-1054]。众所周知,最大似然估计量以及更一般地最小对比度估计量确实存在各种弱点,并且它们的惩罚版本也是如此。特别是,它们并不稳健,并且需要对模型和基本参数集进行限制性假设才能正确工作。我们提出了一种替代的建设,它来自一个合适的度量空间中的一些概率球之间的许多同时测试的估计。在许多情况下,虽然不是在所有情况下,它会导致一个惩罚的M-估计限制到一个合适的可数参数集。一方面,这种结构应被视为一个理论,而不是一个实用的工具,因为它的高计算复杂性。另一方面,它解决了许多前面提到的困难,前提是我们的构建中涉及的测试存在,这是各种统计框架的情况,包括从iid变量的密度估计或估计具有已知方差的高斯序列的均值。对于所有这样的框架,我们的估计的鲁棒性属性允许以统一的方式处理极大极小估计和模型选择,因为限制极大极小风险相当于使用单个精心选择的模型执行我们的方法。对于这些框架,这导致仅基于参数空间的一些度量性质的极大极小风险的简单界限。此外,该方法适用于各种统计框架,可以同时处理基本上所有类型的模型,线性或非线性,参数和非参数。它还提供了一个简单的方法来聚合初步估计。
This paper is devoted to the definition and study of a family of model selection oriented estimators that we shall call T-estimators (“T” for tests). Their construction is based on former ideas about deriving estimators from some families of tests due to Le Cam [LM Le Cam, Convergence of estimates under dimensionality restrictions, Ann. Statist. 1 (1973) 38–53 and LM Le Cam, On local and global properties in the theory of asymptotic normality of experiments, in: M. Puri (Ed.), Stochastic Processes and Related Topics, vol. 1, Academic Press, New York, 1975, pp. 13–54] and Birgé [L. Birgé, Approximation dans les espaces métriques et théorie de l’estimation, Z. Wahrscheinlichkeitstheorie Verw. Gebiete 65 (1983) 181–237, L. Birgé, Sur un théorème de minimax et son application aux tests, Probab. Math. Statist. 3 (1984) 259–282 and L. Birgé, Stabilité et instabilité du risque minimax pour des variables indépendantes équidistribuées, Ann. Inst. H. Poincaré Sect. B 20 (1984) 201–223] and about complexity based model selection from Barron and Cover [AR Barron, TM Cover, Minimum complexity density estimation, IEEE Trans. Inform. Theory 37 (1991) 1034–1054].It is well-known that maximum likelihood estimators and, more generally, minimum contrast estimators do suffer from various weaknesses, and their penalized versions as well. In particular they are not robust and they require restrictive assumptions on both the models and the underlying parameter set to work correctly. We propose an alternative construction, which derives an estimator from many simultaneous tests between some probability balls in a suitable metric space. In many cases, although not in all, it results in a penalized M-estimator restricted to a suitable countable set of parameters. On the one hand, this construction should be considered as a theoretical rather than a practical tool because of its high computational complexity. On the other hand, it solves many of the previously mentioned difficulties provided that the tests involved in our construction exist, which is the case for various statistical frameworks including density estimation from iid variables or estimating the mean of a Gaussian sequence with a known variance. For all such frameworks, the robustness properties of our estimators allow to deal with minimax estimation and model selection in a unified way, since bounding the minimax risk amounts to performing our method with a single, well-chosen, model. This results, for those frameworks, in simple bounds for the minimax risk solely based on some metric properties of the parameter space. Moreover the method applies to various statistical frameworks and can handle essentially all types of models, linear or not, parametric and non-parametric, simultaneously. It also provides a simple way of aggregating preliminary estimators.