Classification with many classes: Challenges and pluses

Classification with many classes: Challenges and pluses
复制标题

DOI:
10.1016/j.jmva.2019.104536
复制
发表时间:
2019-11-01
影响因子:
1.6
通讯作者:
Pensky, Marianna
Pensky, Marianna
中科院分区:
数学2区
文献类型:
--
作者:
Abramovich, Felix;Pensky, Marianna

文献摘要

被引文献

相似文献

本文的目的是研究高维环境下的多类分类精度,其中类的数量也很大(“大L,大p,小n”模型)。虽然这个问题出现在许多实际应用中,并且最近开发了许多技术来解决这个问题,但据我们所知,没有人对这个重要的设置进行严格的理论分析。本文的目的是填补这一空白。我们考虑一个最常见的设置,分类的高维法向量,不像标准的假设,类的数量可能很大。我们推导出非渐近条件的影响,显着的功能,以及成功的功能选择和分类所需的类之间的距离的下限和上限与给定的精度。此外,我们研究了一个渐近设置的类的数量是发散的特征空间的维数,而每个类的样本数可能是有限的。我们指出一个有趣的,乍一看,有点违反直觉的现象,大量的类可能是一个“祝福”,而不是一个“诅咒”,因为在某些设置中,分类的精度可以提高类的数量的增长。这是由于更准确的特征选择,因为即使是较弱的重要特征,其强度不足以在粗略分类中表现出来,在类别之间共享,随着类别数量的增加,具有更强的影响。我们补充我们的理论研究的模拟研究和真实的数据的例子,我们再次观察到上述现象。(C)2019爱思唯尔公司All rights reserved.
The objective of the paper is to study accuracy of multi-class classification in high-dimensional setting, where the number of classes is also large ("large L, large p, small n" model). While this problem arises in many practical applications and many techniques have been recently developed for its solution, to the best of our knowledge nobody provided a rigorous theoretical analysis of this important setup. The purpose of the present paper is to fill in this gap.We consider one of the most common settings, classification of high-dimensional normal vectors where, unlike standard assumptions, the number of classes could be large. We derive non-asymptotic conditions on effects of significant features, and the low and the upper bounds for distances between classes required for successful feature selection and classification with a given accuracy. Furthermore, we study an asymptotic setup where the number of classes is diverging with the dimension of feature space and while the number of samples per class is possibly limited. We point out on an interesting and, at first glance, somewhat counter-intuitive phenomenon that a large number of classes may be a "blessing" rather than a "curse" since, in certain settings, the precision of classification can improve as the number of classes grows. This is due to more accurate feature selection since even weaker significant features, which are not sufficiently strong to be manifested in a coarse classification, being shared across the classes, have a stronger impact as the number of classes increases. We supplement our theoretical investigation by a simulation study and a real data example where we again observe the above phenomenon. (C) 2019 Elsevier Inc. All rights reserved.