Distributed adaptive nearest neighbor classifier: algorithm and theory

Distributed adaptive nearest neighbor classifier: algorithm and theory
复制标题

DOI:
10.1007/s11222-023-10267-7
复制
发表时间:
2021-05
影响因子:
2.2
通讯作者:
Ruiqi Liu;Ganggang Xu;Zuofeng Shang
Ruiqi Liu;Ganggang Xu;Zuofeng Shang
中科院分区:
数学2区
文献类型:
--
作者:
Ruiqi Liu;Ganggang Xu;Zuofeng Shang

文献摘要

被引文献

相似文献

当数据量非常大或物理存储在不同的位置时,分布式最近邻(NN)分类器是一种有吸引力的分类工具。我们提出了一种新的分布式自适应NN分类器的最近邻的数量是一个调谐参数随机选择的数据驱动的标准。在搜索最优参数时,提出了一种提前停止规则,不仅加快了计算速度,而且改善了算法的有限样本性能。研究了分布式自适应神经网络分类器在不同子样本组成下的超风险收敛速度。特别是,我们表明,当子样本的大小是足够大的,所提出的分类器实现了接近最优的收敛速度。所提出的方法的有效性证明,通过模拟研究以及实证应用到现实世界的数据集。
When data is of an extraordinarily large size or physically stored in different locations, the distributed nearest neighbor (NN) classifier is an attractive tool for classification. We propose a novel distributed adaptive NN classifier for which the number of nearest neighbors is a tuning parameter stochastically chosen by a data-driven criterion. An early stopping rule is proposed when searching for the optimal tuning parameter, which not only speeds up the computation but also improves the finite sample performance of the proposed algorithm. Convergence rate of excess risk of the distributed adaptive NN classifier is investigated under various sub-sample size compositions. In particular, we show that when the sub-sample sizes are sufficiently large, the proposed classifier achieves the nearly optimal convergence rate. Effectiveness of the proposed approach is demonstrated through simulation studies as well as an empirical application to a real-world dataset.