An iterative method for classification of binary data

An iterative method for classification of binary data
复制标题

DOI:
10.1093/imaiai/iaaa003
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Denali Molitor;D. Needell
Denali Molitor;D. Needell
中科院分区:
其他
文献类型:
--
作者:
Denali Molitor;D. Needell

文献摘要

相似文献

在当今数据驱动的世界中,从大规模数据中存储、处理和收集见解是主要挑战。为了存储大量的高维数据,通常需要数据压缩,因此,用于分析压缩数据的有效推理方法是必要的。最近设计的一个简单的框架,使用二进制数据进行分类的基础上,我们证明,可以提高这种方法的分类精度,通过迭代应用程序的输出作为输入到下一个应用程序。作为一个侧面的后果,我们表明,原来的框架可以用作数据预处理步骤,以提高其他方法,如支持向量机的性能。对于几个简单的设置,我们展示了获得理论保证的迭代分类方法的准确性的能力。基本分类框架的简单性使其易于理论分析。
In today’s data-driven world, storing, processing and gleaning insights from large-scale data are major challenges. Data compression is often required in order to store large amounts of high-dimensional data, and thus, efficient inference methods for analyzing compressed data are necessary. Building on a recently designed simple framework for classification using binary data, we demonstrate that one can improve classification accuracy of this approach through iterative applications whose output serves as input to the next application. As a side consequence, we show that the original framework can be used as a data preprocessing step to improve the performance of other methods, such as support vector machines. For several simple settings, we showcase the ability to obtain theoretical guarantees for the accuracy of the iterative classification method. The simplicity of the underlying classification framework makes it amenable to theoretical analysis.