An Interior-Point Method for Large-Scale l1-Regularized Logistic Regression

An Interior-Point Method for Large-Scale l1-Regularized Logistic Regression
复制标题

DOI:
10.5555/1314498.1314550
复制
发表时间:
2007-12
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Kwangmoo Koh;Seung-Jean Kim;Stephen P. Boyd
Kwangmoo Koh;Seung-Jean Kim;Stephen P. Boyd
中科院分区:
其他
文献类型:
--
作者:
Kwangmoo Koh;Seung-Jean Kim;Stephen P. Boyd

文献摘要

被引文献

相似文献

带 l1 正则化的逻辑回归已被提出作为分类问题中特征选择的一种有前景的方法。在本文中,我们描述了一种用于解决大规模 l1 正则化逻辑回归问题的有效内点方法。具有多达一千个左右功能和示例的小问题可以在 PC 上在几秒钟内得到解决;具有数万个特征和示例的中型问题可以在数十秒内解决(假设数据有些稀疏)。基本方法的一种变体,即使用预条件共轭梯度法来计算搜索步骤,可以在几分钟内在 PC 上解决具有一百万个特征和示例(例如 20 个新闻组数据集)的非常大的问题。使用热启动技术,可以比独立解决一系列问题更有效地计算整个正则化路径的良好近似值。
Logistic regression with l1 regularization has been proposed as a promising method for feature selection in classification problems. In this paper we describe an efficient interior-point method for solving large-scale l1-regularized logistic regression problems. Small problems with up to a thousand or so features and examples can be solved in seconds on a PC; medium sized problems, with tens of thousands of features and examples, can be solved in tens of seconds (assuming some sparsity in the data). A variation on the basic method, that uses a preconditioned conjugate gradient method to compute the search step, can solve very large problems, with a million features and examples (e.g., the 20 Newsgroups data set), in a few minutes, on a PC. Using warm-start techniques, a good approximation of the entire regularization path can be computed much more efficiently than by solving a family of problems independently.