Classification of a large microarray data set: Algorithm comparison and analysis of drug signatures

Classification of a large microarray data set: Algorithm comparison and analysis of drug signatures
复制标题

DOI:
10.1101/gr.2807605
复制
发表时间:
2005-05-01
期刊:
影响因子:
7
通讯作者:
Jarnagin, K
Jarnagin, K
中科院分区:
生物学1区
文献类型:
--
作者:
Natsoulis, G;El Ghaoui, L;Jarnagin, K

文献摘要

被引文献

相似文献

已经建立了一个大型的基因表达数据库,描述了数百种已批准和已撤销的药物、毒物和生化标准在活体大鼠各器官中的基因表达和生理效应。为了从这个庞大的数据库中获得有用的生物学知识,使用597微阵列数据子集对各种监督分类算法进行了比较。我们的研究表明,基于油支持向量机(svm)和逻辑回归的几种类型的线性分类器可以用于获得易于解释的具有高分类性能的药物特征。这两种方法都可以调整为以短的形式产生药物治疗的分类器,加权基因列表,经分析显示,一些特征基因对分类决策有积极贡献(作为兴趣类的“奖励”),而其他特征基因有消极贡献(作为“惩罚”)。奖励和惩罚基因的结合通过降低假阳性治疗的数量来提高表现。这些算法的结果与特征选择技术相结合,进一步缩短了药物特征的长度,这是开发有用的诊断生物标志物和低成本分析的重要一步。对于同一分类终点,可以生成多个无共同基因的特征。通过比较这些基因表,可以确定某一类生物过程的特征。
A large gene expression database has been produced that characterizes the gene expression and physiological effects of hundreds of approved and withdrawn drugs, toxicants, and biochemical standards in various organs of live rats. Ill order to derive useful biological knowledge from this large database, a variety Of Supervised classification algorithms were compared using a 597-microarray Subset of the data. Our Studies show that several types of linear classifiers based Oil Support Vector Machines (SVMs) and Logistic Regression can be used to derive readily interpretable drug signatures with high classification performance. Both methods can be tuned to produce classifiers of drug treatments in the form of short, weighted gene lists which upon analysis reveal that some of the signature genes have a positive contribution (act as "rewards" for the class-of-interest) while others have a negative contribution (act as "penalties") to the classification decision. The combination of reward and penalty genes enhances performance by keeping the number of false positive treatments low. The results of these algorithms are combined with feature selection techniques that further reduce the length of the drug signatures, an important step towards the development of useful diagnostic biomarkers and low-cost assays. Multiple signatures with no genes in common can be generated for the same classification end-point. Comparison of these gene lists identifies biological processes characteristic of a given class.