A feature selection method based on multiple kernel learning with expression profiles of different types.

A feature selection method based on multiple kernel learning with expression profiles of different types.
复制标题

一种基于多核学习的不同类型表达谱的特征选择方法

DOI:
10.1186/s13040-017-0124-x
复制
发表时间:
2017
期刊:
影响因子:
4.5
通讯作者:
Liang Y
Liang Y
中科院分区:
生物学3区
文献类型:
--
作者:
Du W;Cao Z;Song T;Li Y;Liang Y

文献摘要

被引文献

相似文献

背景随着高通量技术的发展,研究人员可以从多个公共数据库中获取大量不同类型的表达数据。由于这些数据大多样本数量少,特征数成百上千,如何有效、稳健地利用特征选择技术从表情数据中提取信息丰富的特征是一个具有挑战性和关键性的问题。到目前为止,已经提出了大量的特征选择方法,并应用于不同类型的表情数据的分析。结果本文提出了一种基于多核学习的混合特征选择方法(MKL),并对不同类型的表情数据进行了性能评估。首先,利用MKL的优化功能来度量特征与分类样本之间的相关性。在该步骤中,使用迭代梯度下降过程对支持向量机的参数和核置信度进行优化。然后,通过对每个特征的优化函数进行排序,选择一组相关的特征。在此基础上,提出了一种嵌入前向选择的方法来检测相关特征集中的紧凑特征子集。结论我们不仅比较了与其他方法的分类精度,而且比较了不同算法的稳定性、相似性和一致性。对于使用不同性能度量的不同类型的表达数据集,该方法具有令人满意的特征选择能力。
BackgroundWith the development of high-throughput technology, the researchers can acquire large number of expression data with different types from several public databases. Because most of these data have small number of samples and hundreds or thousands features, how to extract informative features from expression data effectively and robustly using feature selection technique is challenging and crucial. So far, a mass of many feature selection approaches have been proposed and applied to analyse expression data of different types. However, most of these methods only are limited to measure the performances on one single type of expression data by accuracy or error rate of classification.ResultsIn this article, we propose a hybrid feature selection method based on Multiple Kernel Learning (MKL) and evaluate the performance on expression datasets of different types. Firstly, the relevance between features and classifying samples is measured by using the optimizing function of MKL. In this step, an iterative gradient descent process is used to perform the optimization both on the parameters of Support Vector Machine (SVM) and kernel confidence. Then, a set of relevant features is selected by sorting the optimizing function of each feature. Furthermore, we apply an embedded scheme of forward selection to detect the compact feature subsets from the relevant feature set.ConclusionsWe not only compare the classification accuracy with other methods, but also compare the stability, similarity and consistency of different algorithms. The proposed method has a satisfactory capability of feature selection for analysing expression datasets of different types using different performance measurements.