Visualization-based cancer microarray data classification analysis

Visualization-based cancer microarray data classification analysis
复制标题

DOI:
10.1093/bioinformatics/btm312
复制
发表时间:
2007-08-15
期刊:
影响因子:
5.8
通讯作者:
Zupan, Blaz
Zupan, Blaz
中科院分区:
生物学3区
文献类型:
--
作者:
Mramor, Minca;Leban, Gregor;Zupan, Blaz

文献摘要

被引文献

相似文献

动机:分析癌症微阵列数据的方法通常面临两个截然不同的挑战:它们推断的模型在对新的组织样本进行分类时需要表现良好,同时提供对隐藏在数据中的模式和基因相互作用的洞察。最新的有监督数据挖掘方法通常只涵盖这些方面中的一个方面,这推动了具有可靠分类性能的预测模型的开发,这些方法可以很容易地与领域专家进行沟通。结果:数据可视化可以为类标签数据的知识发现和分析提供一种很好的方法。我们之前已经开发了一种称为VizRank的方法,它可以根据不同类的数据实例的分离程度对基于点的可视化进行评分和排名。我们在这里扩展VizRank的技术,以发现异常值,评分特征(基因)和执行分类,以及证明所提出的方法很适合癌症微阵列分析。使用VizRank和radviz可视化技术对一组以前发表的癌症微阵列数据集进行处理,我们能够找到简单、可解释的数据预测,这些数据预测只包括一小部分基因,但确实在不同的癌症类型之间进行了明显的区分。我们还报告了我们通过可视化进行分类的方法获得了与最先进的监督数据挖掘技术相媲美的性能。可用性:VizRank和radviz是作为Orange数据挖掘套件(http://www.ailab.si/orange).)的一部分实现的联系人:blaz.zupan@Fri.uni-lj.si补充信息:补充数据可从http://www.ailab.si/supp/bi-cancer.获得
Motivation: Methods for analyzing cancer microarray data often face two distinct challenges: the models they infer need to perform well when classifying new tissue samples while at the same time providing an insight into the patterns and gene interactions hidden in the data. State-of-the-art supervised data mining methods often cover well only one of these aspects, motivating the development of methods where predictive models with a solid classification performance would be easily communicated to the domain expert.Results: Data visualization may provide for an excellent approach to knowledge discovery and analysis of class-labeled data. We have previously developed an approach called VizRank that can score and rank point-based visualizations according to degree of separation of data instances of different class. We here extend VizRank with techniques to uncover outliers, score features ( genes) and perform classification, as well as to demonstrate that the proposed approach is well suited for cancer microarray analysis. Using VizRank and radviz visualization on a set of previously published cancer microarray data sets, we were able to find simple, interpretable data projections that include only a small subset of genes yet do clearly differentiate among different cancer types. We also report that our approach to classification through visualization achieves performance that is comparable to state-of-the-art supervised data mining techniques.Availability: VizRank and radviz are implemented as part of the Orange data mining suite (http://www.ailab.si/orange). Contact: blaz.zupan@fri.uni-lj.siSupplementary information: Supplementary data are available from http://www.ailab.si/supp/bi-cancer.