Being Robust (in High Dimensions) Can Be Practical

Being Robust (in High Dimensions) Can Be Practical
复制标题

DOI:
--
复制
发表时间:
2017-03
期刊:
--
影响因子:
--
通讯作者:
Ilias Diakonikolas;Gautam Kamath;D. Kane;Jerry Li;Ankur Moitra;Alistair Stewart
Ilias Diakonikolas;Gautam Kamath;D. Kane;Jerry Li;Ankur Moitra;Alistair Stewart
中科院分区:
其他
文献类型:
--
作者:
Ilias Diakonikolas;Gautam Kamath;D. Kane;Jerry Li;Ankur Moitra;Alistair Stewart

文献摘要

被引文献

相似文献

在高维度上,强大的估计比在一个维度上更具挑战性:大多数技术会导致棘手的优化问题,或者仅能容忍一小部分错误的估计器。理论计算机科学的最新工作表明,在适当的分布模型中,有可能与多项式时间算法的均值和协方差估算,这些算法可以忍受与维度无关的持续腐败分数。但是,对于高维应用,这些算法的样本和时间复杂性非常大。在这项工作中,我们通过建立最佳的样本复杂性界限,达到对数因素,并提供各种改进,使算法能够容忍更大的损坏部分。最后,我们在合成数据和真实数据上都表明我们的算法具有最先进的性能,并突然使高维鲁棒估计成为现实的可能性。
Robust estimation is much more challenging in high dimensions than it is in one dimension: Most techniques either lead to intractable optimization problems or estimators that can tolerate only a tiny fraction of errors. Recent work in theoretical computer science has shown that, in appropriate distributional models, it is possible to robustly estimate the mean and covariance with polynomial time algorithms that can tolerate a constant fraction of corruptions, independent of the dimension. However, the sample and time complexity of these algorithms is prohibitively large for high-dimensional applications. In this work, we address both of these issues by establishing sample complexity bounds that are optimal, up to logarithmic factors, as well as giving various refinements that allow the algorithms to tolerate a much larger fraction of corruptions. Finally, we show on both synthetic and real data that our algorithms have state-of-the-art performance and suddenly make high-dimensional robust estimation a realistic possibility.