Feature Selection for Nonlinear Regression and its Application to Cancer Research

Feature Selection for Nonlinear Regression and its Application to Cancer Research
复制标题

DOI:
10.1137/1.9781611974010.9
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
Yijun Sun;Jin Yao;S. Goodison
Yijun Sun;Jin Yao;S. Goodison
中科院分区:
其他
文献类型:
--
作者:
Yijun Sun;Jin Yao;S. Goodison

文献摘要

被引文献

相似文献

特征选择是机器学习中的一个基本问题。随着高通量技术的出现,它在广泛的科学学科中变得越来越重要。本文研究高维非线性回归的特征选择问题。这个问题还没有得到很好的解决,在社区中,现有的方法遭受的问题,如局部极小值,简化的模型假设,高计算复杂性和选择的功能不直接相关的学习精度。我们提出了一个新的包装方法,解决了其中的一些问题。首先,我们开发了一种新的方法来估计样本响应和预测误差,然后部署一个特征加权策略,以找到一个特征子空间的预测误差函数最小化。我们将其制定为SVM框架内的优化问题,并使用迭代方法解决它。在每次迭代中,基于梯度下降的方法推导出有效地找到一个解决方案。一个大规模的模拟研究进行了四个合成和九个癌症微阵列数据集,证明了所提出的方法的有效性。
Feature selection is a fundamental problem in machine learning. With the advent of high-throughput technologies, it becomes increasingly important in a wide range of scientific disciplines. In this paper, we consider the problem of feature selection for high-dimensional nonlinear regression. This problem has not yet been well addressed in the community, and existing methods suffer from issues such as local minima, simplified model assumptions, high computational complexity and selected features not directly related to learning accuracy. We propose a new wrapper method that addresses some of these issues. We start by developing a new approach to estimating sample responses and prediction errors, and then deploy a feature weighting strategy to find a feature subspace where a prediction error function is minimized. We formulate it as an optimization problem within the SVM framework and solve it using an iterative approach. In each iteration, a gradient descent based approach is derived to efficiently find a solution. A large-scale simulation study is performed on four synthetic and nine cancer microarray datasets that demonstrates the effectiveness of the proposed method.