A between-Class Overlapping Filter-Based Method for transcriptome Data Analysis

A between-Class Overlapping Filter-Based Method for transcriptome Data Analysis
复制标题

DOI:
10.1142/s0219720012500102
复制
发表时间:
2012-08
影响因子:
1
通讯作者:
Alok Sharma;S. Imoto;S. Miyano
Alok Sharma;S. Imoto;S. Miyano
中科院分区:
生物学4区
文献类型:
--
作者:
Alok Sharma;S. Imoto;S. Miyano

文献摘要

被引文献

相似文献

特征选择算法在识别和发现用于癌症分类的重要基因方面起着至关重要的作用。特征选择算法可以大致分为两大类:基于过滤器的方法和基于包装器的方法。基于滤波器的方法由于其许多优点而在文献中相当流行,包括计算效率、简单的架构以及发现生物学和临床方面的直观简单的手段。然而,这些方法具有局限性,并且所选择的基因的分类准确性不太准确。在本文中,我们提出了一套单变量过滤器为基础的方法,使用类间重叠的标准。所提出的技术进行了比较与许多其他单变量过滤器为基础的方法,使用急性白血病数据集。以下属性已被检查:所选的单个基因和基因子集的分类准确性;冗余检查选定的基因之间使用岭回归和LASSO方法;相似性和敏感性分析;功能分析;和,稳定性分析。一个全面的实验表明,我们提出的技术有前途的结果。使用类间重叠准则的基于单变量滤波器的方法是准确和鲁棒的,具有生物学意义,并且计算效率高且易于实现。因此,它们非常适合生物学和临床发现。
Feature selection algorithms play a crucial role in identifying and discovering important genes for cancer classification. Feature selection algorithms can be broadly categorized into two main groups: filter-based methods and wrapper-based methods. Filter-based methods have been quite popular in the literature due to their many advantages, including computational efficiency, simplistic architecture, and an intuitively simple means of discovering biological and clinical aspects. However, these methods have limitations, and the classification accuracy of the selected genes is less accurate. In this paper, we propose a set of univariate filter-based methods using a between-class overlapping criterion. The proposed techniques have been compared with many other univariate filter-based methods using an acute leukemia dataset. The following properties have been examined: classification accuracy of the selected individual genes and the gene subsets; redundancy check among selected genes using ridge regression and LASSO methods; similarity and sensitivity analyses; functional analysis; and, stability analysis. A comprehensive experiment shows promising results for our proposed techniques. The univariate filter based methods using between-class overlapping criterion are accurate and robust, have biological significance, and are computationally efficient and easy to implement. Therefore, they are well suited for biological and clinical discoveries.