A feature selection approach for identification of signature genes from SAGE data.

A feature selection approach for identification of signature genes from SAGE data.
复制标题

DOI:
10.1186/1471-2105-8-169
复制
发表时间:
2007-05-22
期刊:
影响因子:
3
通讯作者:
Brentani, Helena
Brentani, Helena
中科院分区:
生物学4区
文献类型:
--
作者:
Barrera, Junior;Cesar, Roberto M Jr;Humes, Carlos Jr;Martins, David C Jr;Patrao, Diogo F C;Silva, Paulo J S;Brentani, Helena

文献摘要

相似文献

基因表达谱分析的一个目标是鉴定出能够有力区分不同类型或等级肿瘤的特征基因。利用微阵列技术已经提出了几种基于表达谱的肿瘤分类器。由于微阵列和SAGE技术的概率模型的重要差异,重要的是开发合适的技术来选择特定的基因从SAGE测量。提出了一种基于SAGE数据分析来选择区分不同生物状态的特定基因的新框架。新的框架适用于强基因的识别,分离的生物状态在一个由训练集的基因表达定义的特征空间中的支撑误差。从SAGE测量的概率模型定义的可信度区间被用来识别区分不同状态的基因,在强基因方法选择的所有基因组中具有更高的可靠性。一个分数考虑到可信度和支持的错误值,以考虑基因组的排名提出。使用SAGE数据从神经胶质瘤获得的结果,从而证实了所介绍的方法。代表计数数据的模型,如SAGE,提供了额外的统计信息,允许更强大的分析。由概率模型提供的额外的统计信息被纳入本文所述的方法。所介绍的方法是适合于识别的签名基因,导致一个良好的分离的生物状态,使用SAGE和可适用于其他计数方法,如大规模并行签名测序(MPSS)或最近的合成测序(SBS)技术。所提出的方法确定的一些这样的基因可能是有用的生成分类器。
One goal of gene expression profiling is to identify signature genes that robustly distinguish different types or grades of tumors. Several tumor classifiers based on expression profiling have been proposed using microarray technique. Due to important differences in the probabilistic models of microarray and SAGE technologies, it is important to develop suitable techniques to select specific genes from SAGE measurements. A new framework to select specific genes that distinguish different biological states based on the analysis of SAGE data is proposed. The new framework applies the bolstered error for the identification of strong genes that separate the biological states in a feature space defined by the gene expression of a training set. Credibility intervals defined from a probabilistic model of SAGE measurements are used to identify the genes that distinguish the different states with more reliability among all gene groups selected by the strong genes method. A score taking into account the credibility and the bolstered error values in order to rank the groups of considered genes is proposed. Results obtained using SAGE data from gliomas are presented, thus corroborating the introduced methodology. The model representing counting data, such as SAGE, provides additional statistical information that allows a more robust analysis. The additional statistical information provided by the probabilistic model is incorporated in the methodology described in the paper. The introduced method is suitable to identify signature genes that lead to a good separation of the biological states using SAGE and may be adapted for other counting methods such as Massive Parallel Signature Sequencing (MPSS) or the recent Sequencing-By-Synthesis (SBS) technique. Some of such genes identified by the proposed method may be useful to generate classifiers.