Conditional Sparse Linear Regression

Conditional Sparse Linear Regression
复制标题

条件稀疏线性回归

DOI:
--
复制
发表时间:
2016
期刊:
Information Technology Convergence and Services
影响因子:
--
通讯作者:
Brendan Juba
Brendan Juba
中科院分区:
--
文献类型:
--
作者:
Brendan Juba

文献摘要

被引文献

相似文献

机器学习和统计学通常专注于构建捕获绝大多数数据的模型,可能会忽略一小部分数据作为“噪音”或“离群值”。“相比之下,在这里我们考虑联合识别一个重要的(但可能很小的)群体部分的问题,其中有一个高度稀疏的线性回归拟合,以及线性拟合的系数。我们认为,这些任务是有趣的,因为模型本身可能能够在这种特殊情况下实现更好的预测,但也因为它们可以帮助我们理解的数据。我们给这样的问题的算法下的超级规范,当这个未知的部分人口描述的k-DNF条件和回归拟合是s-稀疏常数k和s。对于回归拟合不太稀疏或使用期望误差时该问题的变体,我们也给出了初步的算法,并强调该问题是未来工作的挑战。
Machine learning and statistics typically focus on building models that capture the vast majority of the data, possibly ignoring a small subset of data as "noise" or "outliers." By contrast, here we consider the problem of jointly identifying a significant (but perhaps small) segment of a population in which there is a highly sparse linear regression fit, together with the coefficients for the linear fit. We contend that such tasks are of interest both because the models themselves may be able to achieve better predictions in such special cases, but also because they may aid our understanding of the data. We give algorithms for such problems under the sup norm, when this unknown segment of the population is described by a k-DNF condition and the regression fit is s-sparse for constant k and s. For the variants of this problem when the regression fit is not so sparse or using expected error, we also give a preliminary algorithm and highlight the question as a challenge for future work.